daemon-sec-cheatsheet

The cheatsheet vault for operators: AD, enumeration, exploitation, priv-esc, web, DFIR
git clone https://git.daemon-sec.xyz/daemon-sec-cheatsheet.git
Log | Files | Refs | README | LICENSE

usernames-and-accounts.md (41967B)


      1 ---
      2 title: "Username & Account Discovery"
      3 description: "Pivot from one handle to every platform it appears on, and work out which hits are actually the same person."
      4 category: osint
      5 subcategory: "People & Identity"
      6 tags: [osint, username, accounts, pivoting]
      7 tools: [sherlock, maigret, blackbird, whatsmyname, naminter, socid-extractor, ghunt, imagehash]
      8 difficulty: beginner
      9 updated: 2026-10-04
     10 references:
     11   - name: "Bellingcat's Online Investigation Toolkit"
     12     url: "https://bellingcat.gitbook.io/toolkit"
     13     author: "Bellingcat"
     14     license: none
     15     relation: derived
     16     note: "Tool catalogue: names, descriptions, cost flags and links for this area."
     17   - name: "OSINT Newsletter Tools Library"
     18     url: "https://tools.osintnewsletter.com"
     19     author: "The OSINT Newsletter"
     20     license: none
     21     relation: derived
     22     note: "Second tool catalogue, cross-checked against the above."
     23   - name: "WhatsMyName"
     24     url: "https://github.com/WebBreacher/WhatsMyName"
     25     author: "Micah Hoffman and contributors"
     26     relation: link-only
     27     note: "CC BY-SA 4.0. The detection-field schema and the dataset counts quoted on this page were computed from wmn-data.json; no description text is reproduced."
     28   - name: "Sherlock"
     29     url: "https://github.com/sherlock-project/sherlock"
     30     author: "Sherlock Project"
     31     relation: link-only
     32     note: "Flags verified against sherlock_project/sherlock.py; site counts computed from its bundled data.json."
     33   - name: "Maigret"
     34     url: "https://github.com/soxoj/maigret"
     35     author: "soxoj"
     36     relation: link-only
     37     note: "Flags, identifier types and report formats verified against maigret/maigret.py."
     38   - name: "Naminter documentation"
     39     url: "https://3xp0rt.github.io/Naminter/"
     40     author: "3xp0rt"
     41     license: MIT
     42     relation: link-only
     43     note: "CLI reference: options, detection modes and result statuses were verified against this."
     44   - name: "Blackbird"
     45     url: "https://github.com/p1ngul1n0/blackbird"
     46     author: "p1ngul1n0"
     47     relation: link-only
     48     note: "Flags verified against blackbird.py."
     49   - name: "socid_extractor"
     50     url: "https://github.com/soxoj/socid-extractor"
     51     author: "soxoj"
     52     relation: link-only
     53     note: "CLI flags verified against socid_extractor/cli.py."
     54   - name: "Chess.com Published-Data API"
     55     url: "https://www.chess.com/news/view/published-data-api"
     56     author: "Chess.com"
     57     relation: link-only
     58     note: "Player-profile endpoint and field names used in the worked example."
     59 ---
     60 
     61 ## What this covers
     62 
     63 You have one username. You want every other place that username exists, and then you want to know
     64 which of those are the same human. The first part is automated and takes four minutes. The second
     65 part is not automated, takes the rest of the day, and is the only part that produces a finding.
     66 
     67 Everything on this page serves the second part. The enumerators are a way of generating candidates
     68 cheaply; treating their output as a result is the single most common failure in this discipline.
     69 
     70 ## Method
     71 
     72 1. **Enumerate the exact handle** across platforms with an automated checker. Fast, noisy, a
     73    starting point only.
     74 2. **Generate variants** before concluding anything. People append birth years, swap separators,
     75    and drop vowels. `j.smith`, `jsmith90`, `j_smith`, `jaysmith` are all the same person often
     76    enough to be worth checking.
     77 3. **Confirm each hit manually.** Enumerators check whether a profile URL resolves. They do not
     78    check whether it is your target. Open it.
     79 4. **Pivot on content, not the handle.** Same profile photo, same bio phrasing, same follower
     80    overlap, same posting hours. A shared handle is weak evidence; a shared photo is strong.
     81 5. **Reach for a stable identifier.** Handles change; the numeric ID behind them usually does not.
     82    Once you have a GAIA ID, a VK id or a GitHub `id`, you have something that survives a rename
     83    and that you can search for independently.
     84 6. **Work the timeline.** Account creation dates that cluster suggest one person registering
     85    everywhere at once. A ten-year-old account and a two-week-old one with the same name are
     86    probably not related.
     87 7. **Write down the odds.** Say how strong each piece of evidence is before you add it up, so that
     88    somebody can argue with a number rather than with your conclusion.
     89 
     90 The judgement calls: decide early whether you are chasing one person or one handle, because a
     91 handle that turns out to be held by nine strangers needs a different write-up. Decide whether the
     92 handle came from the target or from somebody describing the target, since a transcription error
     93 upstream wastes the whole day. And decide what you will accept as confirmation before you start
     94 looking, because the evidence always looks more convincing once you have spent six hours on it.
     95 
     96 ## The match arithmetic
     97 
     98 Two separate calculations sit underneath this work, and conflating them is how people end up
     99 certain about the wrong person. The first is about the tool: how often does an enumerator claim a
    100 profile that is not there. The second is about the person: given that a handle matches, how likely
    101 is it the same human.
    102 
    103 ### What an enumerator actually tests
    104 
    105 A checker does not know what a profile looks like. It has four numbers and strings per site, and
    106 it compares a response against them. In the WhatsMyName schema those fields are `e_code` and
    107 `e_string` for an account that exists, `m_code` and `m_string` for one that does not. Sherlock's
    108 own database expresses the same idea as a single `errorType` of `status_code`, `message` or
    109 `response_url`, with an `errorMsg` or `errorCode` beside it.
    110 
    111 The failure mode falls straight out of the data. Counted from `wmn-data.json` on 4 October 2026:
    112 
    113 ```text
    114 WhatsMyName dataset
    115   sites                                           717   across 20 categories
    116   e_code == m_code (status cannot discriminate)   175   24.4 per cent
    117   m_code == 200    (soft 404 for a missing user)  178
    118   entries carrying a `protection` flag             62   34 of them Cloudflare
    119   POST-based entries                               23
    120 
    121 Sherlock bundled data.json
    122   sites                                           481
    123   errorType = status_code                         327   68.0 per cent
    124   errorType = message                             127
    125   errorType = response_url                         27
    126   entries with a regexCheck (handle-shape filter)  95
    127   entries with a separate urlProbe                 52
    128 ```
    129 
    130 Recompute either of those yourself rather than trusting the numbers above; they move every week:
    131 
    132 ```bash
    133 curl -sL -o wmn-data.json https://raw.githubusercontent.com/WebBreacher/WhatsMyName/main/wmn-data.json
    134 
    135 python3 - <<'PY'
    136 import json, collections
    137 d = json.load(open("wmn-data.json"))
    138 s = d["sites"]
    139 same = [x for x in s if x.get("e_code") == x.get("m_code")]
    140 print("sites           ", len(s))
    141 print("code-ambiguous  ", len(same), f"{100*len(same)/len(s):.1f}%")
    142 print("soft 404 (m=200)", sum(1 for x in s if x.get("m_code") == 200))
    143 print("by category     ", collections.Counter(x["cat"] for x in s).most_common(5))
    144 PY
    145 ```
    146 
    147 So on roughly a quarter of the dataset the HTTP status is worthless and the body string is the
    148 only discriminator. Steam and Keybase are ordinary examples: both return 200 whether or not the
    149 account exists. This is why a checker's **detection mode** is not a cosmetic setting. Naminter
    150 names the two rules explicitly:
    151 
    152 ```text
    153 --mode all   EXISTS only when the status AND the body string both match the "exists" pair.
    154              One of the two matching gives PARTIAL_EXISTS, not EXISTS.           (default)
    155 
    156 --mode any   EXISTS when the status OR the body string matches.
    157 ```
    158 
    159 Under `any`, a status-only rule fires on all 178 soft-404 sites for a handle that exists nowhere
    160 on earth. That is your false-positive budget in one number: switching mode can move the hit count
    161 by up to a quarter of the list without a single new account being involved. Run strict, read the
    162 partials as a separate pile, and never report a count taken from a loose run.
    163 
    164 ### What a handle match is worth
    165 
    166 Now the second calculation. Let `k` be the number of distinct humans who hold the handle across
    167 the platforms you checked. Before any content evidence at all, a candidate account is your target
    168 with probability `1/k`:
    169 
    170 ```text
    171 k =  1   p = 1.00      the handle is yours alone; rare, and worth confirming rather than assuming
    172 k =  2   p = 0.50
    173 k =  3   p = 0.33      odds 1:2 against
    174 k =  5   p = 0.20
    175 k = 40   p = 0.025     a dictionary word or firstname+year; the handle is near-worthless alone
    176 ```
    177 
    178 `k` is not a guess. You estimate it by opening the hits and counting the ones that are plainly
    179 different people — a 2011 account posting in another language, a brand, a bot. Each one you can
    180 positively exclude raises `k` and lowers what the handle is worth.
    181 
    182 Evidence then multiplies the odds:
    183 
    184 ```text
    185 posterior odds = prior odds x LR1 x LR2 x ...
    186 
    187 LR = P(evidence | same person) / P(evidence | different people)
    188 ```
    189 
    190 The likelihood ratios you assign are judgements, and the discipline is to write them down:
    191 
    192 ```text
    193 identical avatar bitmap, image found nowhere else   LR ~ 50-100   strong
    194 identical avatar, image is a stock photo or meme    LR ~ 1        worthless
    195 registration within the same hour on two platforms  LR ~ 10-30    moderate
    196 same rare bio typo or identical punctuation habit   LR ~ 10-50    moderate to strong
    197 same stable platform ID reached from both accounts  not an LR     it is the same account
    198 ```
    199 
    200 The trap is independence. Odds only multiply when the pieces of evidence are conditionally
    201 independent of each other. A matching handle and a matching display name are **not** independent:
    202 both are derived from the same real name, so counting them twice inflates your confidence by a
    203 factor you never earned. Neither are two accounts on platforms that offer "sign in with Google"
    204 and copy the avatar across on your behalf. Before you multiply, ask what common cause could have
    205 produced both observations, and if there is one, use the stronger of the two and discard the
    206 other.
    207 
    208 The last row of that table is the point of the whole page. An inference from an avatar is an
    209 argument. A stable numeric identifier reached from both ends is not an inference at all, and it is
    210 what you should be trying to get to.
    211 
    212 ## Key tools
    213 
    214 ### Sherlock
    215 
    216 The usual first pass. It constructs the expected profile URL for each site in its bundled
    217 `data.json` and reads the response, so it touches only public pages and needs no API keys. 481
    218 sites in the current list, which is a fraction of what Maigret and the WhatsMyName dataset carry —
    219 Sherlock's value is that it is fast, dependency-light and present in most distro repositories.
    220 
    221 ```bash
    222 pipx install sherlock-project
    223 # or: dnf install sherlock-project
    224 # or: docker run -it --rm sherlock/sherlock user123
    225 ```
    226 
    227 ```bash
    228 # single username; results also land in a text file named for the handle
    229 sherlock user123
    230 
    231 # several at once
    232 sherlock alice bob charlie
    233 
    234 # the separator sweep: {?} expands to '_', '-' and '.' in turn,
    235 # so this one invocation covers john_doe, john-doe and john.doe
    236 sherlock 'john{?}doe'
    237 
    238 # narrow to named sites -- the names are the keys in data.json, case-sensitive
    239 sherlock user123 --site GitHub --site Instagram
    240 
    241 # machine-readable output for the triage spreadsheet
    242 sherlock user123 --csv
    243 sherlock user123 --xlsx
    244 sherlock user123 --txt --folderoutput ./runs/kestrel
    245 
    246 # show the misses too, which is how you tell "not found" from "never checked"
    247 sherlock user123 --print-all
    248 
    249 # route through a proxy, since 481 requests from one address in a few seconds is a pattern
    250 sherlock user123 --proxy socks5://127.0.0.1:1080 --timeout 20
    251 
    252 # when a single site's verdict looks wrong, read the response it actually got
    253 sherlock user123 --site Steam --dump-response
    254 
    255 # run against an unmerged upstream site list, or a local one you have edited
    256 sherlock user123 --json https://example.org/data.json
    257 sherlock user123 --local
    258 ```
    259 
    260 Interpreting the output: a green `[+]` is a URL that matched the site's detection rule, nothing
    261 more. Sites excluded upstream for being unreliable are skipped silently unless you pass
    262 `--ignore-exclusions`, which, as its own help text says, brings back false positives along with
    263 the coverage.
    264 
    265 How it misleads you: 327 of the 481 entries decide on the HTTP status code alone, so every
    266 soft-404 site in the list is a coin toss. Ninety-five entries carry a `regexCheck` that rejects
    267 handles which cannot exist on that platform — useful, but it means a handle containing a dot or a
    268 dash will be reported as absent from sites that simply do not permit the character, which is a
    269 different statement from "nobody registered it". And the `--tor` flag that circulates in older
    270 tutorials no longer exists; current releases have `--proxy` only, and passing `--tor` will stop
    271 the run with an argparse error rather than quietly running unproxied.
    272 
    273 ### Maigret
    274 
    275 Broader than Sherlock and considerably more useful after the first pass, because it extracts
    276 profile fields rather than only reporting existence. It also searches by identifiers other than
    277 usernames, which is how you re-enter the search from a numeric ID.
    278 
    279 ```bash
    280 pipx install maigret      # or: pip install maigret, sudo snap install maigret
    281 docker pull soxoj/maigret
    282 ```
    283 
    284 ```bash
    285 # default run: the 500 highest-traffic sites
    286 maigret user123
    287 
    288 # everything in the database, which is where the long-tail regional sites live
    289 maigret user123 -a
    290 
    291 # trade coverage for speed on a first sweep
    292 maigret user123 --top-sites 100
    293 
    294 # scope by site tag; run --stats first to see which tags the database actually uses
    295 maigret user123 --stats
    296 maigret user123 --tags photo,dating
    297 maigret user123 --exclude-tags "xx NSFW xx"
    298 
    299 # reports: -H html, -P pdf, -C csv, -T txt, -M markdown, -G graph, -X xmind
    300 maigret user123 -H -C --folderoutput ./runs/kestrel
    301 
    302 # JSON for a pipeline; TYPE is 'simple' or 'ndjson'
    303 maigret user123 -J ndjson
    304 
    305 # pivot from a profile page: parse it, extract IDs and usernames, search on those
    306 maigret --parse https://www.deviantart.com/muse1908
    307 
    308 # search by a stable identifier instead of a handle
    309 maigret 112233445566778899001 --id-type gaia_id
    310 maigret 12345678 --id-type vk_id
    311 
    312 # combine two name fragments into candidate handles
    313 maigret john smith --permute
    314 
    315 # highlight sites whose page body also mentions a term you care about
    316 maigret user123 --keywords photography lisbon
    317 
    318 # throttle and proxy: 6,000-odd database entries at default concurrency is a very loud burst
    319 maigret user123 -n 20 --timeout 30 --proxy socks5://127.0.0.1:1080
    320 maigret user123 --tor-proxy socks5://127.0.0.1:9050
    321 ```
    322 
    323 The valid `--id-type` values are fixed by the source and worth knowing, because they are the exit
    324 doors from handle-based reasoning: `username`, `yandex_public_id`, `gaia_id`, `vk_id`, `ok_id`,
    325 `wikimapia_uid`, `steam_id`, `uidme_uguid`, `yelp_userid`, `orcid`, `qq_id`, `bilibili_id`.
    326 
    327 Before you trust a long run, measure the list:
    328 
    329 ```bash
    330 # tests each site against the known-claimed and known-unclaimed handles in the database
    331 maigret --self-check --diagnose
    332 
    333 # and disable the ones that fail, so the next run is quieter
    334 maigret --self-check --auto-disable
    335 ```
    336 
    337 How it misleads you: recursive search is on by default, so Maigret will take a username it found
    338 on one profile page and start searching for *that*, which is excellent until the profile belongs
    339 to someone else and the report quietly becomes a report about two people. Pass `--no-recursion`
    340 when you want a clean single-handle run, and read the report's own account-tree rather than the
    341 flat list. The `--ai` summary is generated text about your results, not a finding, and should
    342 never appear in a write-up. Note also that `--print-not-found` is off by default, so a short
    343 report can mean a short list of sites rather than a short list of accounts.
    344 
    345 ### WhatsMyName and Naminter
    346 
    347 [WhatsMyName](https://whatsmyname.app/) is the community-maintained detection dataset — a single
    348 `wmn-data.json` of 717 site entries under CC BY-SA 4.0 — rather than a checker. The project
    349 removed its bundled scripts in 2023 and now maintains only the data, which several other tools
    350 consume. Its detection strings tend to be the most current of any list, because fixing one is a
    351 one-line pull request.
    352 
    353 Three inclusion rules shape what it can ever find, and they are the reason an absence here means
    354 very little: a site must be publicly accessible without a login, the username must appear in the
    355 profile URL, and the site must not swap the username for a numeric ID. Every platform that
    356 identifies people by number is therefore structurally invisible to this entire class of tool.
    357 
    358 [Naminter](https://github.com/3xp0rt/Naminter) is the checker built specifically for that dataset,
    359 and it is the one to use when you care about the difference between a hit and a maybe.
    360 
    361 ```bash
    362 pip install naminter      # or: uv tool install naminter
    363 docker run --rm -it ghcr.io/3xp0rt/naminter --username john_doe
    364 naminter --version
    365 ```
    366 
    367 ```bash
    368 # strict detection is the default: status AND string must both match
    369 naminter -u user123
    370 
    371 # show every status, not just the hits -- this is the run you actually read
    372 naminter -u user123 --filter-all
    373 
    374 # the partials are the pile that needs eyes; the conflicts are a broken site entry
    375 naminter -u user123 --filter-partial --filter-conflicting
    376 
    377 # scope by category; the dataset's 20 values include social, coding, gaming, finance
    378 naminter -u user123 --include-categories social --include-categories coding
    379 naminter -u user123 --exclude-categories "xx NSFW xx"
    380 
    381 # keep the evidence: raw response bodies plus a report, before anything is deleted
    382 naminter -u user123 --filter-all --save-response --response-dir ./evidence \
    383     --csv --csv-path hits.csv --html --html-path report.html
    384 
    385 # slow it down and proxy it
    386 naminter -u user123 --proxy socks5://127.0.0.1:9050 --timeout 60 --max-tasks 20
    387 
    388 # pin the dataset to a local copy so a run is reproducible months later
    389 naminter -u user123 --local-data ./wmn-data.json --local-schema ./wmn-data-schema.json
    390 
    391 # measure the list before you trust it: checks every site against its known accounts
    392 naminter --test
    393 naminter --test -s GitHub -s Reddit -vvv
    394 ```
    395 
    396 Read the status symbols rather than counting lines. `+` exists, `-` missing, `~+` and `~-` are
    397 partial matches where only one of the two criteria fired, `*` is conflicting — the exists and
    398 missing indicators both matched, which means the site entry is stale — `?` unknown, `X` a site
    399 marked invalid in the data, and `!` a request error. A partial is not a weak hit; it is a site
    400 where the tool cannot tell, and the fix is to open the URL.
    401 
    402 How it misleads you: `--impersonate` defaults to `chrome`, so Naminter is deliberately not
    403 truthful about what it is. That is usually what you want, but it means a site serving a different
    404 page to automation may give a verdict you cannot reproduce with `curl`, and `--impersonate none`
    405 will sometimes flip a result. `--max-tasks` defaults to 50, which against 717 sites is a visible
    406 burst from one address. And `--filter-exists` is the default, so a quiet run is not a negative
    407 result until you have re-run it with `--filter-all`.
    408 
    409 ### Blackbird
    410 
    411 Searches by username and by email from the same binary, which is its reason to exist: a single
    412 pass over both identifier types, with permutation built in. Useful as a cross-check when Sherlock
    413 and Naminter disagree, because it is a third implementation reading a different list.
    414 
    415 ```bash
    416 git clone https://github.com/p1ngul1n0/blackbird
    417 cd blackbird
    418 pip install -r requirements.txt
    419 ```
    420 
    421 ```bash
    422 python blackbird.py --username user123
    423 python blackbird.py --email target@example.com
    424 
    425 # several at once, or from a file
    426 python blackbird.py -u alice bob charlie
    427 python blackbird.py -uf handles.txt
    428 python blackbird.py -ef addresses.txt
    429 
    430 # permutation: --permute ignores single elements, --permuteall combines everything
    431 python blackbird.py -u john smith --permute
    432 python blackbird.py -u john michael smith --permuteall
    433 
    434 # scope by a list property
    435 python blackbird.py -u user123 --filter "cat=social"
    436 python blackbird.py -u user123 --no-nsfw
    437 
    438 # keep the pages, not just the verdicts
    439 python blackbird.py -u user123 --dump --csv --json
    440 
    441 # throttle, proxy, and freeze the site list for a reproducible run
    442 python blackbird.py -u user123 --proxy socks5://127.0.0.1:9050 \
    443     --timeout 60 --max-concurrent-requests 10 --no-update
    444 ```
    445 
    446 How it misleads you: it updates its site lists on every run unless you pass `--no-update`, so two
    447 runs a week apart are not comparable and neither is reproducible — pin it with `--no-update` the
    448 moment a run matters. The `--ai` profiling requires `--setup-ai` and an API key, and it produces
    449 characterisation of a person from thin evidence, which is exactly the thing an investigation
    450 should not be doing with a language model. Treat the account list as the output and ignore the
    451 rest.
    452 
    453 ### socid_extractor
    454 
    455 This is the step that turns a list of URLs into evidence. It parses a profile page or API response
    456 and returns a flat dictionary of fields, including the stable internal identifiers a platform uses
    457 behind the handle — GAIA ID for Google, Facebook UID, Yandex public ID, Instagram `pk`. Those
    458 survive renames, which handles do not. It is the extraction engine Maigret uses, and it is worth
    459 running directly on each confirmed hit.
    460 
    461 ```bash
    462 pipx install socid-extractor     # or: pip install socid-extractor
    463 ```
    464 
    465 ```bash
    466 # one profile
    467 socid_extractor --url https://www.deviantart.com/muse1908
    468 
    469 # structured, for the case file
    470 socid_extractor --url https://github.com/torvalds --json
    471 
    472 # a page you already saved, which is the right way to work from archived evidence
    473 socid_extractor --file ./evidence/profile.html --json
    474 
    475 # sites that need a session
    476 socid_extractor --url https://example.com/profile --cookies 'sessionid=...; csrftoken=...'
    477 socid_extractor --cookie-jar ./cookies.txt --url https://example.com/profile
    478 
    479 # batch runs: skip the HTTP request when the URL matches no known scheme
    480 socid_extractor --url https://example.com/foo --skip-fetch-if-no-url-hint
    481 ```
    482 
    483 Output is `key: value` lines, or JSON with `--json`. The fields worth extracting first are any
    484 `*_id`, `created_at` and the avatar URL; the rest is context.
    485 
    486 How it misleads you: it reads what the page serves, so a page behind Cloudflare, a login wall or
    487 an A/B test returns nothing and that is not the same as the account not existing. The field
    488 ontology is normalised across platforms, which means `created_at` from two sites may have been
    489 derived two different ways — one from an API field, one from a rendered string — and comparing
    490 them to the minute is over-reading. The `--ai-fallback` extractor guesses at pages with no
    491 matching scheme; anything it produces is unverified and should be marked as such.
    492 
    493 ### GHunt
    494 
    495 The Google-specific case, included here because a GAIA ID is the most portable stable identifier
    496 you are likely to recover, and because `maigret --id-type gaia_id` takes one directly.
    497 
    498 ```bash
    499 pipx install ghunt
    500 ghunt login            # authenticates GHunt to Google; required before the modules work
    501 ```
    502 
    503 ```bash
    504 ghunt email target@gmail.com --json user_data.json
    505 ghunt gaia 112233445566778899001 --json gaia.json
    506 ```
    507 
    508 The modules are `login`, `email`, `gaia`, `drive`, `geolocate` and `spiderdal`, and `--json` works
    509 with all but `login`. Email-side work — validation, breach and registration checks, recovery-hint
    510 pivots — belongs on [Email & Phone Number OSINT](/sheets/osint/email-and-phone) rather than here.
    511 
    512 How it misleads you: it runs as an authenticated Google session, so the account you log in with is
    513 visible to Google and the results depend on what that account is permitted to see. Use a research
    514 account, never a personal one.
    515 
    516 ### Name and handle variant generation
    517 
    518 Enumeration tests the strings you give it, so the strings are the whole game. Two web tools
    519 generate them, and both are worth a pass before you conclude a handle is absent.
    520 
    521 [Bellingcat's Name Variant Search](https://bellingcat.github.io/name-variant-search/) generates
    522 plausible alternative forms of a person's name — the problem it solves is transliteration and
    523 spelling across alphabets, where a registry search fails not because the person is absent but
    524 because the name was romanised differently. [NAMINT](https://seintpl.github.io/NAMINT/) works the
    525 other way round, taking first, middle and last names and producing username and email
    526 combinations.
    527 
    528 Then there are the mechanical variants, which you can generate without a tool:
    529 
    530 ```text
    531 Given "John Smith", born 1990, known handle jsmith:
    532   separators     jsmith   j_smith   j-smith   j.smith
    533   order          smithj   smith_j   johns     sjohn
    534   year suffixes  jsmith90  jsmith1990  jsmith09
    535   vowel drops    jsmth    jnsmth
    536   leet           j5mith   jsm1th
    537   platform caps  usernames over the length limit get truncated differently per site
    538 ```
    539 
    540 Sherlock's `{?}` covers only the three separators. Maigret's `--permute` and Blackbird's
    541 `--permute` / `--permuteall` combine name fragments you supply. Nothing generates the year
    542 suffixes or the vowel drops for you, so build that list by hand and feed it as multiple
    543 positional usernames.
    544 
    545 How this misleads you: variant generation multiplies your request volume by the number of
    546 variants, and it multiplies your false-positive count by the same factor. Ten variants against 717
    547 sites is 7,170 requests and roughly ten times the soft-404 noise. Generate variants, but run them
    548 scoped to a category or a site list rather than across everything.
    549 
    550 ## Enumerators at a glance
    551 
    552 | Tool | Site list | Detection rule | Extracts profile data | Reports |
    553 | --- | --- | --- | --- | --- |
    554 | [Sherlock](https://github.com/sherlock-project/sherlock) | Own `data.json`, 481 sites | Status code on 327 of 481; message or response URL on the rest | No | csv, xlsx, txt |
    555 | [Maigret](https://github.com/soxoj/maigret) | Own database, top 500 by default, `-a` for all | Per-site, plus recursive search on extracted usernames | Yes, via socid_extractor | html, pdf, csv, txt, md, graph, xmind, json, neo4j |
    556 | [Naminter](https://github.com/3xp0rt/Naminter) | WhatsMyName, 717 sites | Strict `all` by default; exposes partial and conflicting states | No | csv, json, html, pdf, raw responses |
    557 | [Blackbird](https://github.com/p1ngul1n0/blackbird) | Own list, auto-updating | Per-site | Dumps HTML with `--dump` | csv, json, pdf |
    558 | [WhatsMyName web](https://whatsmyname.app/) | WhatsMyName, 717 sites | As the dataset defines | No | csv |
    559 | [socid_extractor](https://github.com/soxoj/socid-extractor) | Not an enumerator; 250+ parse schemes | n/a | Yes, including stable IDs | json |
    560 
    561 Pick on the second and third columns, not the first. Sherlock for a fast look, Naminter when the
    562 count has to be defensible, Maigret when you want the profile fields in the same pass, Blackbird
    563 as the third opinion, socid_extractor on every hit that survives.
    564 
    565 ## Tool reference
    566 
    567 | Tool | What it does | Cost |
    568 | --- | --- | --- |
    569 | [192](http://www.192.com/) | Searching for someone's address in the UK, phone number and who they live with according to electoral rolls. | free |
    570 | [Bellingcat Name Variant Search](https://bellingcat.github.io/name-variant-search/) | Quickly search many variants of a person's name on Google | free |
    571 | [Blackbird](https://github.com/p1ngul1n0/blackbird) | Check usernames and email addresses on websites and social networks | free |
    572 | [Epieos](https://tools.epieos.com/holehe.php) | Checks where an email has been used. Based on Holehe. | paid |
    573 | [Ghunt](https://github.com/mxrch/GHunt) | A command line tool for obtaining information about Google accounts. | free |
    574 | [Maigret](https://github.com/soxoj/maigret) | Maigret is a Python script that retrieves user information by searching for usernames across various websites and social media platforms. | free |
    575 | [NeutrOSINT](https://github.com/Kr0wZ/NeutrOSINT) | A tool for investigating Proton Mail addresses. | free |
    576 | [Sherlock](https://github.com/sherlock-project/sherlock) | Allows a user to search for the presence of specific usernames across more than 400 websites and social networks. | free |
    577 | [WhatsMyName](https://whatsmyname.app/) | Search for usernames on several hundred platforms | free |
    578 
    579 ## Pitfalls
    580 
    581 - **A 200 is not a profile.** 178 of WhatsMyName's 717 entries record HTTP 200 for an account that
    582   does not exist, and 327 of Sherlock's 481 entries decide on the status code alone. On those
    583   sites a hit is a coin toss until you open the URL.
    584 - **`--mode any` is not "more thorough".** A loose detection rule can promote up to a quarter of
    585   the list to EXISTS without a single real account being involved. The extra hits are not weak
    586   evidence; they are not evidence.
    587 - **Partials are the interesting pile, not the discard pile.** A PARTIAL_EXISTS means the status
    588   matched and the string did not, or the reverse, which usually means the site redesigned. Those
    589   are the sites where a real account is most likely to be hiding from the tool.
    590 - **A CONFLICTING result is a bug report, not a finding.** Both the exists and missing indicators
    591   matched, so the site entry is stale. Fix it upstream or ignore the site; do not average the two.
    592 - **The dataset cannot see numeric-ID platforms at all.** WhatsMyName only includes sites where
    593   the username appears in the URL and is not transformed into an internal ID. Absence from every
    594   enumerator you own is therefore consistent with a large presence on platforms that identify
    595   people by number.
    596 - **Handles are recycled.** A handle released by one account and claimed by another means the
    597   2019 archive and the live profile are different people wearing the same name. Check creation
    598   dates before you treat an archived page and a current one as the same account.
    599 - **Display name and handle are not independent evidence.** Both derive from the real name.
    600   Multiplying them together is how a 1:2 prior becomes a confident, wrong answer.
    601 - **A matching avatar is only as good as the image's rarity.** An identical bitmap that also
    602   appears on 4,000 other profiles is worth nothing. Reverse-search the avatar before you score
    603   it — see [Reverse Image Search](/sheets/osint/reverse-image-search).
    604 - **Self-check the list before you trust a long run.** `maigret --self-check --diagnose` and
    605   `naminter --test` both measure the detection rules against known accounts. A list you have not
    606   measured gives you a hit count with no error bar.
    607 - **Browser impersonation changes the answer.** Naminter impersonates Chrome by default and
    608   Blackbird can be made to; a verdict you cannot reproduce with a plain client is a verdict that
    609   depends on a header. Record which mode produced it.
    610 - **Enumerators are loud.** 717 sites at Naminter's default 50 concurrent requests is a burst from
    611   one address in a few seconds. Proxy it, lower `--max-tasks` or `-n`, and assume the target can
    612   see the pattern if they operate any of the sites.
    613 - **Variant sweeps multiply the noise, not just the coverage.** Ten variants is ten times the
    614   requests and ten times the soft-404 hits. Scope them by category.
    615 - **Coverage is skewed.** These lists are heavy on English-language and global platforms and thin
    616   on regional ones. Absence from a tool's results is not absence from the internet.
    617 - **Paid aggregators recycle stale data.** A "verified" result from a people-search service is
    618   often a years-old scrape. Check when the underlying data was collected — see
    619   [People Search & Public Records](/sheets/osint/people-search).
    620 - **The AI summary is not a finding.** Maigret's `--ai` and Blackbird's `--ai` produce generated
    621   prose about your results. It cites nothing, and it will characterise a person from three data
    622   points. Keep it out of the write-up.
    623 - **Archive before you cite.** Profiles are deleted while you are looking at them. Capture with
    624   `--save-response` or `--dump` and preserve properly — see
    625   [Archiving & Evidence Preservation](/sheets/osint/archiving-and-evidence).
    626 
    627 ## Worked example
    628 
    629 One datum: the handle `kestrel_ait`, seen once in a forum signature. No name, no email, no claim
    630 about who it belongs to. The question asked is narrow — does the person behind the GitHub account
    631 with that handle also hold the Chess.com account with that handle.
    632 
    633 1. **Sweep the separators.** Sherlock's `{?}` expands to `_`, `-` and `.`, which covers the three
    634    mechanical variants in one invocation:
    635 
    636 ```bash
    637 sherlock 'kestrel{?}ait' --csv --folderoutput ./runs/kestrel
    638 ```
    639 
    640    `kestrel-ait` and `kestrel.ait` return two hits each, both on soft-404 sites, both empty pages
    641    when opened. The underscore form is the live one and the rest of the work uses it.
    642 
    643 2. **Run the strict pass and read every status.** Naminter against the full WhatsMyName dataset,
    644    with responses preserved:
    645 
    646 ```bash
    647 naminter -u kestrel_ait --filter-all --mode all \
    648     --save-response --response-dir ./evidence \
    649     --csv --csv-path kestrel.csv --max-tasks 20
    650 ```
    651 
    652 ```text
    653 [+] GitHub            https://github.com/kestrel_ait
    654 [+] Chess.com         https://api.chess.com/pub/player/kestrel_ait
    655 [+] Last.fm           https://www.last.fm/user/kestrel_ait
    656 [+] Mastodon (infosec.exchange)
    657 [+] SoundCloud        https://soundcloud.com/kestrel_ait
    658 [+] Patreon           https://www.patreon.com/kestrel_ait
    659 [+] Gravatar          https://en.gravatar.com/kestrel_ait.json
    660 [+] Steam             https://steamcommunity.com/id/kestrel_ait
    661 [+] Keybase           https://keybase.io/_/api/1.0/user/lookup.json?usernames=kestrel_ait
    662 
    663 exists 9   partial 23   conflicting 4   error 6   missing 675   total 717
    664 ```
    665 
    666 3. **Quantify the ambiguity rather than guessing at it.** The same run with the loose rule:
    667 
    668 ```bash
    669 naminter -u kestrel_ait --filter-exists --mode any --max-tasks 20
    670 ```
    671 
    672    returns **36** hits. The 27 extra are exactly the 23 partials plus the 4 conflicts, promoted by
    673    a rule that accepts a status match on its own. Nothing was discovered between the two runs. If
    674    the report had been written from the `any` run it would have claimed 36 accounts, of which at
    675    most 9 were ever candidates.
    676 
    677 4. **Open all nine.** Steam and Keybase are the soft-404 sites from the arithmetic above: both
    678    return 200 and both pages are empty. Gravatar resolves to a default identicon with no profile.
    679    Patreon is a reserved-but-unused page. That leaves five: GitHub, Chess.com, Last.fm, SoundCloud
    680    and the Mastodon account.
    681 
    682 5. **Estimate `k`.** Last.fm has scrobbles since 2014, all of one genre, with a listed location in
    683    another country. SoundCloud is a label account with a logo avatar and a contact address on a
    684    company domain. Both are positively excluded. So at least three distinct humans or entities
    685    hold this handle: the GitHub owner, the Last.fm owner, the SoundCloud owner.
    686 
    687 ```text
    688 k = 3   ->   prior that Chess.com is the GitHub person = 1/3 = 0.33   (odds 1:2 against)
    689 ```
    690 
    691 6. **Pull the stable identifiers.** Both remaining platforms publish an API:
    692 
    693 ```bash
    694 socid_extractor --url https://github.com/kestrel_ait --json
    695 curl -s https://api.chess.com/pub/player/kestrel_ait
    696 ```
    697 
    698 ```json
    699 {
    700   "player_id": 196341287,
    701   "@id": "https://api.chess.com/pub/player/kestrel_ait",
    702   "username": "kestrel_ait",
    703   "status": "basic",
    704   "avatar": "https://images.chesscomfiles.com/uploads/v1/user/196341287.9f1c2a7e.200x200o.png",
    705   "joined": 1696153260,
    706   "last_online": 1759512000,
    707   "followers": 4,
    708   "country": "https://api.chess.com/pub/country/GB"
    709 }
    710 ```
    711 
    712    GitHub returns `"id": 146902331` and `"created_at": "2023-10-01T10:12:04Z"`. Neither ID is the
    713    other, and no identifier is shared, so there is no account-level link. This stays an inference.
    714 
    715 7. **Do the timestamp arithmetic properly.** `joined` is a Unix epoch and has to be converted
    716    before it means anything:
    717 
    718 ```bash
    719 python3 -c "import datetime as dt; print(dt.datetime.fromtimestamp(1696153260, dt.timezone.utc))"
    720 # 2023-10-01 09:41:00+00:00
    721 ```
    722 
    723 ```text
    724 Chess.com joined   2023-10-01 09:41:00 UTC
    725 GitHub created_at  2023-10-01 10:12:04 UTC
    726 gap                31 minutes
    727 ```
    728 
    729 8. **Score the avatar, and check its rarity first.** Both avatars are a cropped photograph of a
    730    bird of prey. Reverse-image searching it returns no other instance, so it is not a stock image:
    731 
    732 ```python
    733 from PIL import Image
    734 import imagehash
    735 
    736 gh = imagehash.phash(Image.open("evidence/github_avatar.png"))
    737 cc = imagehash.phash(Image.open("evidence/chesscom_avatar.png"))
    738 print(gh, cc, gh - cc)    # c3a1e4b2d8f09156 c3a1e4b2d8f09156 0
    739 ```
    740 
    741    A 64-bit pHash distance of 0 on a non-stock image is the strongest single piece of evidence in
    742    the file.
    743 
    744 9. **Multiply, with the numbers written down.**
    745 
    746 ```text
    747 prior odds                                   1 : 2
    748 avatar, distance 0, image found nowhere else   x 60
    749 registration 31 minutes apart                  x 20
    750                                              ---------
    751 posterior odds                              1200 : 2  =  600 : 1     p = 0.998
    752 ```
    753 
    754    Those two likelihood ratios are judgements, not measurements, and they are stated so that
    755    somebody can halve them. Even halved twice the answer does not change, which is the useful
    756    property of writing them down. The display name was deliberately not counted: both accounts
    757    read "K. Aitken", and that is the same evidence as the handle, not a second piece.
    758 
    759 What you can assert: two accounts, GitHub `id` 146902331 and Chess.com `player_id` 196341287,
    760 registered 31 minutes apart on 1 October 2023 and carrying a byte-identical avatar that appears
    761 nowhere else, are the same person with posterior odds of roughly 600:1 given a prior of 1:2 from
    762 three observed holders of the handle. No shared platform identifier was recovered, so this is an
    763 inference from independent evidence and not an account-level link, and the write-up should say so
    764 in those words.
    765 
    766 What would falsify it: a fourth and fifth holder of the handle, which would lower the prior; any
    767 instance of that avatar image elsewhere on the internet, which would collapse its likelihood ratio
    768 from 60 to about 1 and take the conclusion with it; or a Chess.com `joined` timestamp read in the
    769 wrong zone — the epoch is UTC, and treating it as local time would move the gap from 31 minutes to
    770 an hour or more and remove the clustering argument entirely. The measurement that carries the most
    771 risk is the avatar's rarity, not its hash distance. The distance is exact and tolerates nothing
    772 above about 2 before it stops meaning "same file"; the rarity is a search result that can be wrong
    773 in one direction only, and it is the one claim here that must be re-checked before publication.
    774 
    775 ## Broader catalogues
    776 
    777 - [Username OSINT](https://tools.osintnewsletter.com/tool-categories/username-osint)
    778 - [People OSINT](https://tools.osintnewsletter.com/tool-categories/people-osint)
    779 
    780 
    781 ## More tools
    782 
    783 Further tools for this area from the OSINT Newsletter Tools Library ([Username OSINT](https://tools.osintnewsletter.com/tool-categories/username-osint)), excluding those already listed above.
    784 
    785 | Tool | What it does |
    786 | --- | --- |
    787 | [Analyst Research Tools](https://analystresearchtools.com/) | A browser-based investigative research platform pulling together a collection of free lookup tools to help you gather, organise… |
    788 | [DeepFind.Me Username Search](https://deepfind.me/tools/social-media/username-search) | Username search tool that scans multiple online platforms to identify where a specific username appears across social media… |
    789 | [Hippie OSINT Toolkit](https://osint.hippie.cat/) | An OSINT web toolkit that allows you to reverse search a domain, a TikTok post, an image, or username (and more). |
    790 | [LoLArchiver](https://lolarchiver.com/) | A niche OSINT tool designed to archive and surface user-related data across gaming and online platforms like League of Legends… |
    791 | [NAMINT](https://seintpl.github.io/NAMINT/) | A web app that generates multiple username, email, and naming combinations from a person’s first, middle, and last name. |
    792 | [OSINT Industries](https://app.osint.industries/) | An all-encompassing OSINT platform that gathers and correlates publicly available digital data such as emails, domains, phone… |
    793 | [Osintly](https://osint.ly/) | A unified OSINT platform that searches for information across pseudonyms, email addresses, phone numbers, IP addresses, domains… |
    794 | [PSNProfiles](https://psnprofiles.com/) | A PlayStation trophy tracking and player statistics platform that allows users to view public PlayStation Network profiles… |
    795 | [Revealer.US username lookup](https://revealer.us/) | An exposure-intelligence and OSINT tool that lets you search for information connected to an email address, username, or phone… |
    796 | [Sherlock Eye](https://www.sherlockeye.io/) | AI-powered OSINT assistant that automates investigations, links data, and surfaces insights fast. |
    797 | [TheBigBrother](https://github.com/chadi0x/TheBigBrother) | All-in-one OSINT toolkit that helps track usernames across the web, investigate emails and domains, extract metadata, and follow… |
    798 | [Threat Actor Username Search](https://threatactorusernames.com/) | A simple OSINT tool that checks whether a username has been observed on known cybercriminal forums and underground platforms. |
    799 
    800 ## Sources
    801 
    802 Both catalogues below are maintained by other people and are considerably larger than
    803 this page. Use them as the canonical index; this sheet is a working route through them.
    804 
    805 - [Bellingcat's Online Investigation Toolkit](https://bellingcat.gitbook.io/toolkit) — ~340 tools, each with its own
    806   review page covering cost, difficulty, requirements and limitations.
    807 - [OSINT Newsletter Tools Library](https://tools.osintnewsletter.com) — ~280 tools, organised by investigative goal.
    808 
    809 Neither publishes a licence, so nothing here is copied from them: tool names, one-line
    810 descriptions, cost flags and links are catalogue facts, and the method and commentary are
    811 this site's own. See [credits](/credits).