daemon-sec-cheatsheet

The cheatsheet vault for operators: AD, enumeration, exploitation, priv-esc, web, DFIR
git clone https://git.daemon-sec.xyz/daemon-sec-cheatsheet.git
Log | Files | Refs | README | LICENSE

index.md (6560B)


      1 ---
      2 title: "Encoding and Transformations"
      3 topic: "Encoding Transformations"
      4 topicSlug: "encoding-transformations"
      5 sourcePath: "Encoding Transformations/README.md"
      6 sourceUrl: "https://github.com/swisskyrepo/PayloadsAllTheThings/blob/3ac27901c711/Encoding%20Transformations/README.md"
      7 sha: "3ac27901c711"
      8 isReadme: true
      9 ---
     10 
     11 # Encoding and Transformations
     12 
     13 > Encoding and Transformations are techniques that change how data is represented or transferred without altering its core meaning. Common examples include URL encoding, Base64, HTML entity encoding, and Unicode transformations. Attackers use these methods as gadgets to bypass input filters, evade web application firewalls, or break out of sanitization routines.
     14 
     15 ## Summary
     16 
     17 * [Unicode](#unicode)
     18     * [Unicode Normalization](#unicode-normalization)
     19     * [Punycode](#punycode)
     20 * [Base64](#base64)
     21 * [Labs](#labs)
     22 * [References](#references)
     23 
     24 ## Unicode
     25 
     26 Unicode is a universal character encoding standard used to represent text from virtually every writing system in the world. Each character (letters, numbers, symbols, emojis) is assigned a unique code point (for example, U+0041 for "A"). Unicode encoding formats like UTF-8 and UTF-16 specify how these code points are stored as bytes.
     27 
     28 ### Unicode Normalization
     29 
     30 Unicode normalization is the process of converting Unicode text into a standardized, consistent form so that equivalent characters are represented the same way in memory.
     31 
     32 [Unicode Normalization reference table](https://appcheck-ng.com/wp-content/uploads/unicode_normalization.html)
     33 
     34 * **NFC** (Normalization Form Canonical Composition): Combines decomposed sequences into precomposed characters where possible.
     35 * **NFD** (Normalization Form Canonical Decomposition): Breaks characters into their decomposed forms (base + combining marks).
     36 * **NFKC** (Normalization Form Compatibility Composition): Like NFC, but also replaces characters with compatibility equivalents (may change appearance/format).
     37 * **NFKD** (Normalization Form Compatibility Decomposition): Like NFD, but also decomposes compatibility characters.
     38 
     39 | Character     | Payload               | After Normalization   |
     40 | ------------- | --------------------- | --------------------- |
     41 | `‥` (U+2025)  | `‥/‥/‥/etc/passwd`    | `../../../etc/passwd` |
     42 | `︰` (U+FE30) | `︰/︰/︰/etc/passwd` | `../../../etc/passwd` |
     43 | `'` (U+FF07) | `' or '1'='1`     | `' or '1'='1`         |
     44 | `"` (U+FF02) | `" or "1"="1`     | `" or "1"="1`         |
     45 | `﹣` (U+FE63) | `admin'﹣﹣`          | `admin'--`            |
     46 | `。` (U+3002) | `domain。com`         | `domain.com`          |
     47 | `/` (U+FF0F) | `//domain.com`      | `//domain.com`        |
     48 | `<` (U+FF1C) | `<img src=a>`       | `<img src=a/>`        |
     49 | `﹛` (U+FE5B) | `﹛﹛3+3﹜﹜`         | `{{3+3}}`             |
     50 | `[` (U+FF3B) | `[[5+5]]`         | `[[5+5]]`             |
     51 | `&` (U+FF06) | `&&whoami`          | `&&whoami`            |
     52 | `p` (U+FF50) | `shell.pʰp`         | `shell.php`           |
     53 | `ʰ` (U+02B0)  | `shell.pʰp`         | `shell.php`           |
     54 | `ª` (U+00AA)  | `ªdmin`               | `admin`               |
     55 
     56 ```py
     57 import unicodedata
     58 string = "ᴾᵃʸˡᵒᵃᵈˢ𝓐𝓵𝓵𝕋𝕙𝕖𝒯𝒽𝒾𝓃ℊ𝓈"
     59 print ('NFC: ' + unicodedata.normalize('NFC', string))
     60 print ('NFD: ' + unicodedata.normalize('NFD', string))
     61 print ('NFKC: ' + unicodedata.normalize('NFKC', string))
     62 print ('NFKD: ' + unicodedata.normalize('NFKD', string))
     63 ```
     64 
     65 ### Punycode
     66 
     67 Punycode is a way to represent Unicode characters (including non-ASCII letters, symbols, and scripts) using only the limited set of ASCII characters (letters, digits, and hyphens).
     68 
     69 It's mainly used in the Domain Name System (DNS), which traditionally supports only ASCII. Punycode allows internationalized domain names (IDNs), so that domain names can include characters from many languages by converting them into a safe ASCII form.
     70 
     71 | Visible in Browser (IDN support) | Actual ASCII (Punycode) |
     72 | -------------------------------- | ----------------------- |
     73 | раypal.com                       | xn--ypal-43d9g.com      |
     74 | paypal.com                       | paypal.com              |
     75 
     76 In MySQL, similar character are treated as equal. This behavior can be abused in Password Reset, Forgot Password, and OAuth Provider sections.
     77 
     78 ```sql
     79 SELECT 'a' = 'ᵃ';
     80 +-------------+
     81 | 'a' = 'ᵃ'   |
     82 +-------------+
     83 |           1 |
     84 +-------------+
     85 ```
     86 
     87 This trick works the SQL query uses `COLLATE utf8mb4_0900_as_cs`.
     88 
     89 ```sql
     90 SELECT 'a' = 'ᵃ' COLLATE utf8mb4_0900_as_cs;
     91 +----------------------------------------+
     92 | 'a' = 'ᵃ' COLLATE utf8mb4_0900_as_cs   |
     93 +----------------------------------------+
     94 |                                      0 |
     95 +----------------------------------------+
     96 ```
     97 
     98 ## Base64
     99 
    100 Base64 encoding is a method for converting binary data (like images or files) or text with special characters into a readable string that uses only ASCII characters (A-Z, a-z, 0-9, +, and /). Every 3 bytes of input are divided into 4 groups of 6 bits and mapped to 4 Base64 characters. If the input isn't a multiple of 3 bytes, the output is padded with `=` characters.
    101 
    102 ```ps1
    103 echo -n admin | base64                            
    104 YWRtaW4=
    105 
    106 echo -n YWRtaW4= | base64 -d
    107 admin
    108 ```
    109 
    110 ## Labs
    111 
    112 * [NahamCon - Puny-Code: 0-Click Account Takeover](https://github.com/VoorivexTeam/white-box-challenges/tree/main/punycode)
    113 * [PentesterLab - Unicode and NFKC](https://pentesterlab.com/exercises/unicode-transform)
    114 
    115 ## References
    116 
    117 * [Puny-Code, 0-Click Account Takeover - Voorivex - June 1, 2025](https://web.archive.org/web/20251211233427/https://blog.voorivex.team/puny-code-0-click-account-takeover)
    118 * [Unicode normalization vulnerabilities - Lazar - September 30, 2021](https://web.archive.org/web/20251224043224/https://lazarv.com/posts/unicode-normalization-vulnerabilities/)
    119 * [Unicode Normalization Vulnerabilities & the Special K Polyglot - AppCheck - September 2, 2019](https://web.archive.org/web/20190916002602/https://appcheck-ng.com/unicode-normalization-vulnerabilities-the-special-k-polyglot/)
    120 * [WAF Bypassing with Unicode Compatibility - Jorge Lajara - February 19, 2020](https://web.archive.org/web/20251230185141/https://jlajara.gitlab.io/Bypass_WAF_Unicode)
    121 * [When "Zoë" !== "Zoë". Or why you need to normalize Unicode strings - Alessandro Segala - March 11, 2019](https://web.archive.org/web/20260128220322/https://withblue.ink/2019/03/11/why-you-need-to-normalize-unicode-strings.html)