daemon-sec-cheatsheet

The cheatsheet vault for operators: AD, enumeration, exploitation, priv-esc, web, DFIR
git clone https://git.daemon-sec.xyz/daemon-sec-cheatsheet.git
Log | Files | Refs | README | LICENSE

unicode-normalization.md (13453B)


      1 ---
      2 title: "Unicode Normalization"
      3 section: "Web Pentesting"
      4 sectionSlug: "pentesting-web"
      5 sourcePath: "src/pentesting-web/unicode-injection/unicode-normalization.md"
      6 sourceUrl: "https://github.com/HackTricks-wiki/hacktricks/blob/188de82beb54e70956b2952367a0af91d26758b8/src/pentesting-web/unicode-injection/unicode-normalization.md"
      7 sha: "188de82beb54e70956b2952367a0af91d26758b8"
      8 isIndex: false
      9 modified: true
     10 license: "CC-BY-NC-4.0"
     11 ---
     12 
     13 # Unicode Normalization
     14 
     15 **This is a summary of:** [**https://appcheck-ng.com/unicode-normalization-vulnerabilities-the-special-k-polyglot/**](https://appcheck-ng.com/unicode-normalization-vulnerabilities-the-special-k-polyglot/). Check it for further details (images taken from there).<sup>[[1]](#references)</sup>
     16 
     17 ## Understanding Unicode and Normalization
     18 
     19 Unicode normalization is a process that ensures different binary representations of characters are standardized to the same binary value. This process is crucial in dealing with strings in programming and data processing. The Unicode standard defines two types of character equivalence:
     20 
     21 1. **Canonical Equivalence**: Characters are considered canonically equivalent if they have the same appearance and meaning when printed or displayed.
     22 2. **Compatibility Equivalence**: A weaker form of equivalence where characters may represent the same abstract character but can be displayed differently.
     23 
     24 There are **four Unicode normalization algorithms**: NFC, NFD, NFKC, and NFKD. Each algorithm employs canonical and compatibility normalization techniques differently. For a more in-depth understanding, you can explore these techniques on [Unicode.org](https://unicode.org/).
     25 
     26 For offensive testing, **NFKC/NFKD** are usually the most interesting forms because they can fold **compatibility characters** (fullwidth symbols, superscripts, modifier letters, ligatures, etc.) into plain ASCII. Also, don't mix this with **homoglyph/confusable** detection: two strings can look identical to a human and still **not** normalize to the same bytes. For phishing/IDN lookalikes check [Homograph / Homoglyph Attacks](https://github.com/HackTricks-wiki/hacktricks/blob/188de82beb54e70956b2952367a0af91d26758b8/src/generic-methodologies-and-resources/phishing-methodology/homograph-attacks.md).<sup>[[9]](#references)</sup>
     27 
     28 ### Key Points on Unicode Encoding
     29 
     30 Understanding Unicode encoding is pivotal, especially when dealing with interoperability issues among different systems or languages. Here are the main points:
     31 
     32 - **Code Points and Characters**: In Unicode, each character or symbol is assigned a numerical value known as a "code point".
     33 - **Bytes Representation**: The code point (or character) is represented by one or more bytes in memory. For instance, LATIN-1 characters (common in English-speaking countries) are represented using one byte. However, languages with a larger set of characters need more bytes for representation.
     34 - **Encoding**: This term refers to how characters are transformed into a series of bytes. UTF-8 is a prevalent encoding standard where ASCII characters are represented using one byte, and up to four bytes for other characters.
     35 - **Processing Data**: Systems processing data must be aware of the encoding used to correctly convert the byte stream into characters.
     36 - **Variants of UTF**: Besides UTF-8, there are other encoding standards like UTF-16 (using a minimum of 2 bytes, up to 4) and UTF-32 (using 4 bytes for all characters).
     37 
     38 It's crucial to comprehend these concepts to effectively handle and mitigate potential issues arising from Unicode's complexity and its various encoding methods.
     39 
     40 An example of how Unicode normalizes two different byte sequences representing the same character:
     41 
     42 ```python
     43 unicodedata.normalize("NFKD","chloe\u0301") == unicodedata.normalize("NFKD", "chlo\u00e9")
     44 ```
     45 
     46 **A list of Unicode equivalent characters can be found here:** [https://appcheck-ng.com/wp-content/uploads/unicode_normalization.html](https://appcheck-ng.com/wp-content/uploads/unicode_normalization.html) and [https://0xacb.com/normalization_table](https://0xacb.com/normalization_table)
     47 
     48 ### Discovering
     49 
     50 If you can find inside a webapp a value that is being echoed back, you could try to send **‘KELVIN SIGN’ (U+0212A)** which **normalises to "K"** (you can send it as `%e2%84%aa`). **If a "K" is echoed back**, then, some kind of **Unicode normalisation** is being performed.
     51 
     52 Another **example**: `%F0%9D%95%83%E2%85%87%F0%9D%99%A4%F0%9D%93%83%E2%85%88%F0%9D%94%B0%F0%9D%94%A5%F0%9D%99%96%F0%9D%93%83` after **unicode** is `Leonishan`.
     53 
     54 A quick way to locally triage candidate characters before sending them to a target is:
     55 
     56 ```python
     57 import unicodedata, urllib.parse
     58 
     59 for c in ["\u212A", "\uFF07", "\uFF02", "\uFF0F", "\uFE64", "\uFE65", "\u00AD"]:
     60     print(hex(ord(c)), urllib.parse.quote(c), unicodedata.normalize("NFKC", c))
     61 ```
     62 
     63 ## **Vulnerable Examples**
     64 
     65 ### **SQL Injection filter bypass**
     66 
     67 Imagine a web page that is using the character `'` to create SQL queries with the user input. This web, as a security measure, **deletes** all occurrences of the character **`'`** from the user input, but **after that deletion** and **before the creation** of the query, it **normalises** using **Unicode** the input of the user.
     68 
     69 Then, a malicious user could insert a different Unicode character equivalent to `' (0x27)` like `%ef%bc%87` , when the input gets normalised, a single quote is created and a **SQLInjection vulnerability** appears:<sup>[[1]](#references)</sup>
     70 
     71 ![https://appcheck-ng.com/unicode-normalization-vulnerabilities-the-special-k-polyglot/](https://raw.githubusercontent.com/HackTricks-wiki/hacktricks/188de82beb54e70956b2952367a0af91d26758b8/src/images/image%20%28702%29.png)<sup>[[1]](#references)</sup>
     72 
     73 **Some interesting Unicode characters**
     74 
     75 - `o` -- %e1%b4%bc
     76 - `r` -- %e1%b4%bf
     77 - `1` -- %c2%b9
     78 - `=` -- %e2%81%bc
     79 - `/` -- %ef%bc%8f
     80 - `-`-- %ef%b9%a3
     81 - `#`-- %ef%b9%9f
     82 - `*`-- %ef%b9%a1
     83 - `'` -- %ef%bc%87
     84 - `"` -- %ef%bc%82
     85 - `|` -- %ef%bd%9c
     86 
     87 ```text
     88 ' or 1=1-- -
     89 %ef%bc%87+%e1%b4%bc%e1%b4%bf+%c2%b9%e2%81%bc%c2%b9%ef%b9%a3%ef%b9%a3+%ef%b9%a3
     90 
     91 " or 1=1-- -
     92 %ef%bc%82+%e1%b4%bc%e1%b4%bf+%c2%b9%e2%81%bc%c2%b9%ef%b9%a3%ef%b9%a3+%ef%b9%a3
     93 
     94 ' || 1==1//
     95 %ef%bc%87+%ef%bd%9c%ef%bd%9c+%c2%b9%e2%81%bc%e2%81%bc%c2%b9%ef%bc%8f%ef%bc%8f
     96 
     97 " || 1==1//
     98 %ef%bc%82+%ef%bd%9c%ef%bd%9c+%c2%b9%e2%81%bc%e2%81%bc%c2%b9%ef%bc%8f%ef%bc%8f
     99 ```
    100 
    101 #### sqlmap template
    102 
    103 
    104 [Sqlmap To Unicode Template](https%3A//github.com/carlospolop/sqlmap_to_unicode_template)
    105 
    106 ### XSS (Cross Site Scripting)
    107 
    108 You could use one of the following characters to trick the webapp and exploit a XSS:<sup>[[1]](#references)</sup>
    109 
    110 ![https://appcheck-ng.com/unicode-normalization-vulnerabilities-the-special-k-polyglot/](https://raw.githubusercontent.com/HackTricks-wiki/hacktricks/188de82beb54e70956b2952367a0af91d26758b8/src/images/image%20%28312%29%20%282%29.png)<sup>[[1]](#references)</sup>
    111 
    112 Notice that for example the first Unicode character proposed can be sent as: `%e2%89%ae` or as `%u226e`
    113 
    114 ![https://appcheck-ng.com/unicode-normalization-vulnerabilities-the-special-k-polyglot/](https://raw.githubusercontent.com/HackTricks-wiki/hacktricks/188de82beb54e70956b2952367a0af91d26758b8/src/images/image%20%28215%29%20%281%29%20%281%29.png)<sup>[[1]](#references)</sup>
    115 
    116 ### Account / identifier collisions
    117 
    118 A very common real-world pattern is **not** direct SQLi/XSS but **cross-endpoint normalization mismatches**:
    119 
    120 1. **Signup / invite** accepts raw Unicode and stores it
    121 2. **Login / password reset / admin search / SSO callback** canonicalizes with NFC/NFKC, `casefold()`, IDNA or a custom transliteration library
    122 3. The uniqueness check, lookup or session binding hits the **wrong account**
    123 
    124 When testing accounts, replay the **same logical identifier** through every flow using raw, lowercase, `casefold()`, NFC, NFKC, punycode and compatibility characters. Useful canaries are `K`, fullwidth forms, combining marks and soft hyphen. If one endpoint reflects/stores the raw value but another one matches a canonicalized version, you may have duplicate-account creation, password-reset confusion or zero-interaction ATO conditions.<sup>[[2]](#references)[[7]](#references)</sup>
    125 
    126 ### Fuzzing Regexes
    127 
    128 When the backend is **checking user input with a regex**, it might be possible that the **input** is being **normalized** for the **regex** but **not** for where it's being **used**. For example, in an Open Redirect or SSRF the regex might be **normalizing the sent URL** but then **accessing it as is**.
    129 
    130 The tool [**recollapse**](https://github.com/0xacb/recollapse) allows to **generate variations of the input** to fuzz the backend. Especially interesting modes are:
    131 
    132 - `3`: normalization
    133 - `6`: case folding / upper / lower
    134 - `7`: byte truncation
    135 
    136 ```bash
    137 recollapse -m 3,6,7 -e 1 'https://legit.example.com'
    138 echo '<svg/onload=alert(1)>' | recollapse | ffuf -w - -u 'https://target/?q=FUZZ' -mc all
    139 ```
    140 
    141 For more info check the [**github repo**](https://github.com/0xacb/recollapse) and this [**post**](https://0xacb.com/2022/11/21/recollapse/).<sup>[[3]](#references)</sup>
    142 
    143 ### Cookie / prefix confusion with Unicode whitespace
    144 
    145 The same bug class also appears in **cookie parsing**: the browser stores/sends one cookie name, but the server later **trims or normalizes** it before applying security rules. If you can inject cookies from a subdomain, response splitting, or XSS, try **leading Unicode whitespace** before `__Host-` / `__Secure-` names and check whether the backend strips it before parsing or validation.
    146 
    147 ```javascript
    148 document.cookie = `${String.fromCodePoint(0x2000)}__Host-session=fixated; Domain=.example.com; Path=/; Secure`;
    149 document.cookie = `${String.fromCodePoint(0x00A0)}__Secure-session=fixated; Domain=.example.com; Path=/; Secure`;
    150 ```
    151 
    152 This is highly implementation-dependent, but recent research showed that some server-side frameworks trim a surprisingly wide range of Unicode whitespace while browsers do not all enforce prefix rules the same way. Treat this as a **parser discrepancy / canonicalization** test whenever cookies are security-critical.<sup>[[4]](#references)</sup>
    153 
    154 ### Best-Fit / WorstFit translations on Windows
    155 
    156 This is **not** standard Unicode normalization, but the exploitation pattern is identical: a security control sees a **safe Unicode character**, and a later Unicode-to-ANSI conversion turns it into dangerous ASCII. On Windows, `U+00AD` (**soft hyphen**) can become `-` on code pages such as **932 / 936 / 950**, which recently led to practical **argument injection** and can also affect **path traversal**, **argument splitting**, and **environment-variable confusion** when ANSI APIs are reached later.<sup>[[5]](#references)</sup>
    157 
    158 Practical probes:
    159 
    160 - Prefix option-like inputs with `%AD` and look for `-` in logs, error messages or downstream process arguments
    161 - Re-test Windows-backed sinks that eventually hit CLI wrappers, archive extractors, CGI binaries or narrow-char filesystem APIs
    162 - Combine `%AD` with other compatibility characters such as fullwidth slash `%EF%BC%8F` when the sink reaches path handling
    163 
    164 ## Unicode Overflow
    165 
    166 From this [blog](https://portswigger.net/research/bypassing-character-blocklists-with-unicode-overflows), the maximum value of a byte is 255, if the server is vulnerable, an overflow can be crafted to produce a specific and unexpected ASCII character. For example, the following characters will be converted to `A`:<sup>[[6]](#references)</sup>
    167 
    168 - 0x4e41
    169 - 0x4f41
    170 - 0x5041
    171 - 0x5141
    172 
    173 This is especially interesting when a sink stores a Unicode value in a **single byte** or otherwise truncates code points before validation/execution. The same research also showed a useful client-side variant with JavaScript `fromCharCode()` overflow:
    174 
    175 ```javascript
    176 String.fromCharCode(0x10000 + 0x31, 0x10000 + 0x33, 0x10000 + 0x33, 0x10000 + 0x37) // 1337
    177 ```
    178 
    179 PortSwigger also added checks/helpers for this class in **ActiveScan++**, **Hackvertor**, and the **Shazzer** Unicode table, so it is worth automating whenever you find brittle ASCII blocklists.<sup>[[6]](#references)</sup>
    180 
    181 ## References
    182 
    183 - [1] [AppCheck - Unicode Normalization Vulnerabilities: The Special-K Polyglot](https://appcheck-ng.com/unicode-normalization-vulnerabilities-the-special-k-polyglot/)
    184 - [2] [Spotify Engineering - Creative usernames](https://engineering.atspotify.com/2013/06/18/creative-usernames/)
    185 - [3] [recollapse - regex fuzzing for normalization discrepancies](https://0xacb.com/2022/11/21/recollapse/)
    186 - [4] [PortSwigger Research - Cookie Chaos: how to bypass Host and Secure cookie prefixes](https://portswigger.net/research/cookie-chaos-how-to-bypass-host-and-secure-cookie-prefixes)
    187 - [5] [Orange Tsai / DEVCORE - WorstFit: Unveiling Hidden Transformers in Windows ANSI!](https://devco.re/blog/2025/01/09/worstfit-unveiling-hidden-transformers-in-windows-ansi/)
    188 - [6] [PortSwigger Research - Bypassing character blocklists with Unicode overflows](https://portswigger.net/research/bypassing-character-blocklists-with-unicode-overflows)
    189 - [7] [Creative usernames and Spotify account hijacking](https://labs.spotify.com/2013/06/18/creative-usernames/)
    190 - [8] [Why does directory traversal attack "%c0%af" work?](https://security.stackexchange.com/questions/48879/why-does-directory-traversal-attack-c0af-work)
    191 - [9] [Bypass WAF with Unicode normalization](https://jlajara.gitlab.io/posts/2020/02/19/Bypass_WAF_Unicode.html)