unicode-normalization.md (13453B)
1 --- 2 title: "Unicode Normalization" 3 section: "Web Pentesting" 4 sectionSlug: "pentesting-web" 5 sourcePath: "src/pentesting-web/unicode-injection/unicode-normalization.md" 6 sourceUrl: "https://github.com/HackTricks-wiki/hacktricks/blob/188de82beb54e70956b2952367a0af91d26758b8/src/pentesting-web/unicode-injection/unicode-normalization.md" 7 sha: "188de82beb54e70956b2952367a0af91d26758b8" 8 isIndex: false 9 modified: true 10 license: "CC-BY-NC-4.0" 11 --- 12 13 # Unicode Normalization 14 15 **This is a summary of:** [**https://appcheck-ng.com/unicode-normalization-vulnerabilities-the-special-k-polyglot/**](https://appcheck-ng.com/unicode-normalization-vulnerabilities-the-special-k-polyglot/). Check it for further details (images taken from there).<sup>[[1]](#references)</sup> 16 17 ## Understanding Unicode and Normalization 18 19 Unicode normalization is a process that ensures different binary representations of characters are standardized to the same binary value. This process is crucial in dealing with strings in programming and data processing. The Unicode standard defines two types of character equivalence: 20 21 1. **Canonical Equivalence**: Characters are considered canonically equivalent if they have the same appearance and meaning when printed or displayed. 22 2. **Compatibility Equivalence**: A weaker form of equivalence where characters may represent the same abstract character but can be displayed differently. 23 24 There are **four Unicode normalization algorithms**: NFC, NFD, NFKC, and NFKD. Each algorithm employs canonical and compatibility normalization techniques differently. For a more in-depth understanding, you can explore these techniques on [Unicode.org](https://unicode.org/). 25 26 For offensive testing, **NFKC/NFKD** are usually the most interesting forms because they can fold **compatibility characters** (fullwidth symbols, superscripts, modifier letters, ligatures, etc.) into plain ASCII. Also, don't mix this with **homoglyph/confusable** detection: two strings can look identical to a human and still **not** normalize to the same bytes. For phishing/IDN lookalikes check [Homograph / Homoglyph Attacks](https://github.com/HackTricks-wiki/hacktricks/blob/188de82beb54e70956b2952367a0af91d26758b8/src/generic-methodologies-and-resources/phishing-methodology/homograph-attacks.md).<sup>[[9]](#references)</sup> 27 28 ### Key Points on Unicode Encoding 29 30 Understanding Unicode encoding is pivotal, especially when dealing with interoperability issues among different systems or languages. Here are the main points: 31 32 - **Code Points and Characters**: In Unicode, each character or symbol is assigned a numerical value known as a "code point". 33 - **Bytes Representation**: The code point (or character) is represented by one or more bytes in memory. For instance, LATIN-1 characters (common in English-speaking countries) are represented using one byte. However, languages with a larger set of characters need more bytes for representation. 34 - **Encoding**: This term refers to how characters are transformed into a series of bytes. UTF-8 is a prevalent encoding standard where ASCII characters are represented using one byte, and up to four bytes for other characters. 35 - **Processing Data**: Systems processing data must be aware of the encoding used to correctly convert the byte stream into characters. 36 - **Variants of UTF**: Besides UTF-8, there are other encoding standards like UTF-16 (using a minimum of 2 bytes, up to 4) and UTF-32 (using 4 bytes for all characters). 37 38 It's crucial to comprehend these concepts to effectively handle and mitigate potential issues arising from Unicode's complexity and its various encoding methods. 39 40 An example of how Unicode normalizes two different byte sequences representing the same character: 41 42 ```python 43 unicodedata.normalize("NFKD","chloe\u0301") == unicodedata.normalize("NFKD", "chlo\u00e9") 44 ``` 45 46 **A list of Unicode equivalent characters can be found here:** [https://appcheck-ng.com/wp-content/uploads/unicode_normalization.html](https://appcheck-ng.com/wp-content/uploads/unicode_normalization.html) and [https://0xacb.com/normalization_table](https://0xacb.com/normalization_table) 47 48 ### Discovering 49 50 If you can find inside a webapp a value that is being echoed back, you could try to send **‘KELVIN SIGN’ (U+0212A)** which **normalises to "K"** (you can send it as `%e2%84%aa`). **If a "K" is echoed back**, then, some kind of **Unicode normalisation** is being performed. 51 52 Another **example**: `%F0%9D%95%83%E2%85%87%F0%9D%99%A4%F0%9D%93%83%E2%85%88%F0%9D%94%B0%F0%9D%94%A5%F0%9D%99%96%F0%9D%93%83` after **unicode** is `Leonishan`. 53 54 A quick way to locally triage candidate characters before sending them to a target is: 55 56 ```python 57 import unicodedata, urllib.parse 58 59 for c in ["\u212A", "\uFF07", "\uFF02", "\uFF0F", "\uFE64", "\uFE65", "\u00AD"]: 60 print(hex(ord(c)), urllib.parse.quote(c), unicodedata.normalize("NFKC", c)) 61 ``` 62 63 ## **Vulnerable Examples** 64 65 ### **SQL Injection filter bypass** 66 67 Imagine a web page that is using the character `'` to create SQL queries with the user input. This web, as a security measure, **deletes** all occurrences of the character **`'`** from the user input, but **after that deletion** and **before the creation** of the query, it **normalises** using **Unicode** the input of the user. 68 69 Then, a malicious user could insert a different Unicode character equivalent to `' (0x27)` like `%ef%bc%87` , when the input gets normalised, a single quote is created and a **SQLInjection vulnerability** appears:<sup>[[1]](#references)</sup> 70 71 <sup>[[1]](#references)</sup> 72 73 **Some interesting Unicode characters** 74 75 - `o` -- %e1%b4%bc 76 - `r` -- %e1%b4%bf 77 - `1` -- %c2%b9 78 - `=` -- %e2%81%bc 79 - `/` -- %ef%bc%8f 80 - `-`-- %ef%b9%a3 81 - `#`-- %ef%b9%9f 82 - `*`-- %ef%b9%a1 83 - `'` -- %ef%bc%87 84 - `"` -- %ef%bc%82 85 - `|` -- %ef%bd%9c 86 87 ```text 88 ' or 1=1-- - 89 %ef%bc%87+%e1%b4%bc%e1%b4%bf+%c2%b9%e2%81%bc%c2%b9%ef%b9%a3%ef%b9%a3+%ef%b9%a3 90 91 " or 1=1-- - 92 %ef%bc%82+%e1%b4%bc%e1%b4%bf+%c2%b9%e2%81%bc%c2%b9%ef%b9%a3%ef%b9%a3+%ef%b9%a3 93 94 ' || 1==1// 95 %ef%bc%87+%ef%bd%9c%ef%bd%9c+%c2%b9%e2%81%bc%e2%81%bc%c2%b9%ef%bc%8f%ef%bc%8f 96 97 " || 1==1// 98 %ef%bc%82+%ef%bd%9c%ef%bd%9c+%c2%b9%e2%81%bc%e2%81%bc%c2%b9%ef%bc%8f%ef%bc%8f 99 ``` 100 101 #### sqlmap template 102 103 104 [Sqlmap To Unicode Template](https%3A//github.com/carlospolop/sqlmap_to_unicode_template) 105 106 ### XSS (Cross Site Scripting) 107 108 You could use one of the following characters to trick the webapp and exploit a XSS:<sup>[[1]](#references)</sup> 109 110 <sup>[[1]](#references)</sup> 111 112 Notice that for example the first Unicode character proposed can be sent as: `%e2%89%ae` or as `%u226e` 113 114 <sup>[[1]](#references)</sup> 115 116 ### Account / identifier collisions 117 118 A very common real-world pattern is **not** direct SQLi/XSS but **cross-endpoint normalization mismatches**: 119 120 1. **Signup / invite** accepts raw Unicode and stores it 121 2. **Login / password reset / admin search / SSO callback** canonicalizes with NFC/NFKC, `casefold()`, IDNA or a custom transliteration library 122 3. The uniqueness check, lookup or session binding hits the **wrong account** 123 124 When testing accounts, replay the **same logical identifier** through every flow using raw, lowercase, `casefold()`, NFC, NFKC, punycode and compatibility characters. Useful canaries are `K`, fullwidth forms, combining marks and soft hyphen. If one endpoint reflects/stores the raw value but another one matches a canonicalized version, you may have duplicate-account creation, password-reset confusion or zero-interaction ATO conditions.<sup>[[2]](#references)[[7]](#references)</sup> 125 126 ### Fuzzing Regexes 127 128 When the backend is **checking user input with a regex**, it might be possible that the **input** is being **normalized** for the **regex** but **not** for where it's being **used**. For example, in an Open Redirect or SSRF the regex might be **normalizing the sent URL** but then **accessing it as is**. 129 130 The tool [**recollapse**](https://github.com/0xacb/recollapse) allows to **generate variations of the input** to fuzz the backend. Especially interesting modes are: 131 132 - `3`: normalization 133 - `6`: case folding / upper / lower 134 - `7`: byte truncation 135 136 ```bash 137 recollapse -m 3,6,7 -e 1 'https://legit.example.com' 138 echo '<svg/onload=alert(1)>' | recollapse | ffuf -w - -u 'https://target/?q=FUZZ' -mc all 139 ``` 140 141 For more info check the [**github repo**](https://github.com/0xacb/recollapse) and this [**post**](https://0xacb.com/2022/11/21/recollapse/).<sup>[[3]](#references)</sup> 142 143 ### Cookie / prefix confusion with Unicode whitespace 144 145 The same bug class also appears in **cookie parsing**: the browser stores/sends one cookie name, but the server later **trims or normalizes** it before applying security rules. If you can inject cookies from a subdomain, response splitting, or XSS, try **leading Unicode whitespace** before `__Host-` / `__Secure-` names and check whether the backend strips it before parsing or validation. 146 147 ```javascript 148 document.cookie = `${String.fromCodePoint(0x2000)}__Host-session=fixated; Domain=.example.com; Path=/; Secure`; 149 document.cookie = `${String.fromCodePoint(0x00A0)}__Secure-session=fixated; Domain=.example.com; Path=/; Secure`; 150 ``` 151 152 This is highly implementation-dependent, but recent research showed that some server-side frameworks trim a surprisingly wide range of Unicode whitespace while browsers do not all enforce prefix rules the same way. Treat this as a **parser discrepancy / canonicalization** test whenever cookies are security-critical.<sup>[[4]](#references)</sup> 153 154 ### Best-Fit / WorstFit translations on Windows 155 156 This is **not** standard Unicode normalization, but the exploitation pattern is identical: a security control sees a **safe Unicode character**, and a later Unicode-to-ANSI conversion turns it into dangerous ASCII. On Windows, `U+00AD` (**soft hyphen**) can become `-` on code pages such as **932 / 936 / 950**, which recently led to practical **argument injection** and can also affect **path traversal**, **argument splitting**, and **environment-variable confusion** when ANSI APIs are reached later.<sup>[[5]](#references)</sup> 157 158 Practical probes: 159 160 - Prefix option-like inputs with `%AD` and look for `-` in logs, error messages or downstream process arguments 161 - Re-test Windows-backed sinks that eventually hit CLI wrappers, archive extractors, CGI binaries or narrow-char filesystem APIs 162 - Combine `%AD` with other compatibility characters such as fullwidth slash `%EF%BC%8F` when the sink reaches path handling 163 164 ## Unicode Overflow 165 166 From this [blog](https://portswigger.net/research/bypassing-character-blocklists-with-unicode-overflows), the maximum value of a byte is 255, if the server is vulnerable, an overflow can be crafted to produce a specific and unexpected ASCII character. For example, the following characters will be converted to `A`:<sup>[[6]](#references)</sup> 167 168 - 0x4e41 169 - 0x4f41 170 - 0x5041 171 - 0x5141 172 173 This is especially interesting when a sink stores a Unicode value in a **single byte** or otherwise truncates code points before validation/execution. The same research also showed a useful client-side variant with JavaScript `fromCharCode()` overflow: 174 175 ```javascript 176 String.fromCharCode(0x10000 + 0x31, 0x10000 + 0x33, 0x10000 + 0x33, 0x10000 + 0x37) // 1337 177 ``` 178 179 PortSwigger also added checks/helpers for this class in **ActiveScan++**, **Hackvertor**, and the **Shazzer** Unicode table, so it is worth automating whenever you find brittle ASCII blocklists.<sup>[[6]](#references)</sup> 180 181 ## References 182 183 - [1] [AppCheck - Unicode Normalization Vulnerabilities: The Special-K Polyglot](https://appcheck-ng.com/unicode-normalization-vulnerabilities-the-special-k-polyglot/) 184 - [2] [Spotify Engineering - Creative usernames](https://engineering.atspotify.com/2013/06/18/creative-usernames/) 185 - [3] [recollapse - regex fuzzing for normalization discrepancies](https://0xacb.com/2022/11/21/recollapse/) 186 - [4] [PortSwigger Research - Cookie Chaos: how to bypass Host and Secure cookie prefixes](https://portswigger.net/research/cookie-chaos-how-to-bypass-host-and-secure-cookie-prefixes) 187 - [5] [Orange Tsai / DEVCORE - WorstFit: Unveiling Hidden Transformers in Windows ANSI!](https://devco.re/blog/2025/01/09/worstfit-unveiling-hidden-transformers-in-windows-ansi/) 188 - [6] [PortSwigger Research - Bypassing character blocklists with Unicode overflows](https://portswigger.net/research/bypassing-character-blocklists-with-unicode-overflows) 189 - [7] [Creative usernames and Spotify account hijacking](https://labs.spotify.com/2013/06/18/creative-usernames/) 190 - [8] [Why does directory traversal attack "%c0%af" work?](https://security.stackexchange.com/questions/48879/why-does-directory-traversal-attack-c0af-work) 191 - [9] [Bypass WAF with Unicode normalization](https://jlajara.gitlab.io/posts/2020/02/19/Bypass_WAF_Unicode.html)