usernames-and-accounts.md (41967B)
1 --- 2 title: "Username & Account Discovery" 3 description: "Pivot from one handle to every platform it appears on, and work out which hits are actually the same person." 4 category: osint 5 subcategory: "People & Identity" 6 tags: [osint, username, accounts, pivoting] 7 tools: [sherlock, maigret, blackbird, whatsmyname, naminter, socid-extractor, ghunt, imagehash] 8 difficulty: beginner 9 updated: 2026-10-04 10 references: 11 - name: "Bellingcat's Online Investigation Toolkit" 12 url: "https://bellingcat.gitbook.io/toolkit" 13 author: "Bellingcat" 14 license: none 15 relation: derived 16 note: "Tool catalogue: names, descriptions, cost flags and links for this area." 17 - name: "OSINT Newsletter Tools Library" 18 url: "https://tools.osintnewsletter.com" 19 author: "The OSINT Newsletter" 20 license: none 21 relation: derived 22 note: "Second tool catalogue, cross-checked against the above." 23 - name: "WhatsMyName" 24 url: "https://github.com/WebBreacher/WhatsMyName" 25 author: "Micah Hoffman and contributors" 26 relation: link-only 27 note: "CC BY-SA 4.0. The detection-field schema and the dataset counts quoted on this page were computed from wmn-data.json; no description text is reproduced." 28 - name: "Sherlock" 29 url: "https://github.com/sherlock-project/sherlock" 30 author: "Sherlock Project" 31 relation: link-only 32 note: "Flags verified against sherlock_project/sherlock.py; site counts computed from its bundled data.json." 33 - name: "Maigret" 34 url: "https://github.com/soxoj/maigret" 35 author: "soxoj" 36 relation: link-only 37 note: "Flags, identifier types and report formats verified against maigret/maigret.py." 38 - name: "Naminter documentation" 39 url: "https://3xp0rt.github.io/Naminter/" 40 author: "3xp0rt" 41 license: MIT 42 relation: link-only 43 note: "CLI reference: options, detection modes and result statuses were verified against this." 44 - name: "Blackbird" 45 url: "https://github.com/p1ngul1n0/blackbird" 46 author: "p1ngul1n0" 47 relation: link-only 48 note: "Flags verified against blackbird.py." 49 - name: "socid_extractor" 50 url: "https://github.com/soxoj/socid-extractor" 51 author: "soxoj" 52 relation: link-only 53 note: "CLI flags verified against socid_extractor/cli.py." 54 - name: "Chess.com Published-Data API" 55 url: "https://www.chess.com/news/view/published-data-api" 56 author: "Chess.com" 57 relation: link-only 58 note: "Player-profile endpoint and field names used in the worked example." 59 --- 60 61 ## What this covers 62 63 You have one username. You want every other place that username exists, and then you want to know 64 which of those are the same human. The first part is automated and takes four minutes. The second 65 part is not automated, takes the rest of the day, and is the only part that produces a finding. 66 67 Everything on this page serves the second part. The enumerators are a way of generating candidates 68 cheaply; treating their output as a result is the single most common failure in this discipline. 69 70 ## Method 71 72 1. **Enumerate the exact handle** across platforms with an automated checker. Fast, noisy, a 73 starting point only. 74 2. **Generate variants** before concluding anything. People append birth years, swap separators, 75 and drop vowels. `j.smith`, `jsmith90`, `j_smith`, `jaysmith` are all the same person often 76 enough to be worth checking. 77 3. **Confirm each hit manually.** Enumerators check whether a profile URL resolves. They do not 78 check whether it is your target. Open it. 79 4. **Pivot on content, not the handle.** Same profile photo, same bio phrasing, same follower 80 overlap, same posting hours. A shared handle is weak evidence; a shared photo is strong. 81 5. **Reach for a stable identifier.** Handles change; the numeric ID behind them usually does not. 82 Once you have a GAIA ID, a VK id or a GitHub `id`, you have something that survives a rename 83 and that you can search for independently. 84 6. **Work the timeline.** Account creation dates that cluster suggest one person registering 85 everywhere at once. A ten-year-old account and a two-week-old one with the same name are 86 probably not related. 87 7. **Write down the odds.** Say how strong each piece of evidence is before you add it up, so that 88 somebody can argue with a number rather than with your conclusion. 89 90 The judgement calls: decide early whether you are chasing one person or one handle, because a 91 handle that turns out to be held by nine strangers needs a different write-up. Decide whether the 92 handle came from the target or from somebody describing the target, since a transcription error 93 upstream wastes the whole day. And decide what you will accept as confirmation before you start 94 looking, because the evidence always looks more convincing once you have spent six hours on it. 95 96 ## The match arithmetic 97 98 Two separate calculations sit underneath this work, and conflating them is how people end up 99 certain about the wrong person. The first is about the tool: how often does an enumerator claim a 100 profile that is not there. The second is about the person: given that a handle matches, how likely 101 is it the same human. 102 103 ### What an enumerator actually tests 104 105 A checker does not know what a profile looks like. It has four numbers and strings per site, and 106 it compares a response against them. In the WhatsMyName schema those fields are `e_code` and 107 `e_string` for an account that exists, `m_code` and `m_string` for one that does not. Sherlock's 108 own database expresses the same idea as a single `errorType` of `status_code`, `message` or 109 `response_url`, with an `errorMsg` or `errorCode` beside it. 110 111 The failure mode falls straight out of the data. Counted from `wmn-data.json` on 4 October 2026: 112 113 ```text 114 WhatsMyName dataset 115 sites 717 across 20 categories 116 e_code == m_code (status cannot discriminate) 175 24.4 per cent 117 m_code == 200 (soft 404 for a missing user) 178 118 entries carrying a `protection` flag 62 34 of them Cloudflare 119 POST-based entries 23 120 121 Sherlock bundled data.json 122 sites 481 123 errorType = status_code 327 68.0 per cent 124 errorType = message 127 125 errorType = response_url 27 126 entries with a regexCheck (handle-shape filter) 95 127 entries with a separate urlProbe 52 128 ``` 129 130 Recompute either of those yourself rather than trusting the numbers above; they move every week: 131 132 ```bash 133 curl -sL -o wmn-data.json https://raw.githubusercontent.com/WebBreacher/WhatsMyName/main/wmn-data.json 134 135 python3 - <<'PY' 136 import json, collections 137 d = json.load(open("wmn-data.json")) 138 s = d["sites"] 139 same = [x for x in s if x.get("e_code") == x.get("m_code")] 140 print("sites ", len(s)) 141 print("code-ambiguous ", len(same), f"{100*len(same)/len(s):.1f}%") 142 print("soft 404 (m=200)", sum(1 for x in s if x.get("m_code") == 200)) 143 print("by category ", collections.Counter(x["cat"] for x in s).most_common(5)) 144 PY 145 ``` 146 147 So on roughly a quarter of the dataset the HTTP status is worthless and the body string is the 148 only discriminator. Steam and Keybase are ordinary examples: both return 200 whether or not the 149 account exists. This is why a checker's **detection mode** is not a cosmetic setting. Naminter 150 names the two rules explicitly: 151 152 ```text 153 --mode all EXISTS only when the status AND the body string both match the "exists" pair. 154 One of the two matching gives PARTIAL_EXISTS, not EXISTS. (default) 155 156 --mode any EXISTS when the status OR the body string matches. 157 ``` 158 159 Under `any`, a status-only rule fires on all 178 soft-404 sites for a handle that exists nowhere 160 on earth. That is your false-positive budget in one number: switching mode can move the hit count 161 by up to a quarter of the list without a single new account being involved. Run strict, read the 162 partials as a separate pile, and never report a count taken from a loose run. 163 164 ### What a handle match is worth 165 166 Now the second calculation. Let `k` be the number of distinct humans who hold the handle across 167 the platforms you checked. Before any content evidence at all, a candidate account is your target 168 with probability `1/k`: 169 170 ```text 171 k = 1 p = 1.00 the handle is yours alone; rare, and worth confirming rather than assuming 172 k = 2 p = 0.50 173 k = 3 p = 0.33 odds 1:2 against 174 k = 5 p = 0.20 175 k = 40 p = 0.025 a dictionary word or firstname+year; the handle is near-worthless alone 176 ``` 177 178 `k` is not a guess. You estimate it by opening the hits and counting the ones that are plainly 179 different people — a 2011 account posting in another language, a brand, a bot. Each one you can 180 positively exclude raises `k` and lowers what the handle is worth. 181 182 Evidence then multiplies the odds: 183 184 ```text 185 posterior odds = prior odds x LR1 x LR2 x ... 186 187 LR = P(evidence | same person) / P(evidence | different people) 188 ``` 189 190 The likelihood ratios you assign are judgements, and the discipline is to write them down: 191 192 ```text 193 identical avatar bitmap, image found nowhere else LR ~ 50-100 strong 194 identical avatar, image is a stock photo or meme LR ~ 1 worthless 195 registration within the same hour on two platforms LR ~ 10-30 moderate 196 same rare bio typo or identical punctuation habit LR ~ 10-50 moderate to strong 197 same stable platform ID reached from both accounts not an LR it is the same account 198 ``` 199 200 The trap is independence. Odds only multiply when the pieces of evidence are conditionally 201 independent of each other. A matching handle and a matching display name are **not** independent: 202 both are derived from the same real name, so counting them twice inflates your confidence by a 203 factor you never earned. Neither are two accounts on platforms that offer "sign in with Google" 204 and copy the avatar across on your behalf. Before you multiply, ask what common cause could have 205 produced both observations, and if there is one, use the stronger of the two and discard the 206 other. 207 208 The last row of that table is the point of the whole page. An inference from an avatar is an 209 argument. A stable numeric identifier reached from both ends is not an inference at all, and it is 210 what you should be trying to get to. 211 212 ## Key tools 213 214 ### Sherlock 215 216 The usual first pass. It constructs the expected profile URL for each site in its bundled 217 `data.json` and reads the response, so it touches only public pages and needs no API keys. 481 218 sites in the current list, which is a fraction of what Maigret and the WhatsMyName dataset carry — 219 Sherlock's value is that it is fast, dependency-light and present in most distro repositories. 220 221 ```bash 222 pipx install sherlock-project 223 # or: dnf install sherlock-project 224 # or: docker run -it --rm sherlock/sherlock user123 225 ``` 226 227 ```bash 228 # single username; results also land in a text file named for the handle 229 sherlock user123 230 231 # several at once 232 sherlock alice bob charlie 233 234 # the separator sweep: {?} expands to '_', '-' and '.' in turn, 235 # so this one invocation covers john_doe, john-doe and john.doe 236 sherlock 'john{?}doe' 237 238 # narrow to named sites -- the names are the keys in data.json, case-sensitive 239 sherlock user123 --site GitHub --site Instagram 240 241 # machine-readable output for the triage spreadsheet 242 sherlock user123 --csv 243 sherlock user123 --xlsx 244 sherlock user123 --txt --folderoutput ./runs/kestrel 245 246 # show the misses too, which is how you tell "not found" from "never checked" 247 sherlock user123 --print-all 248 249 # route through a proxy, since 481 requests from one address in a few seconds is a pattern 250 sherlock user123 --proxy socks5://127.0.0.1:1080 --timeout 20 251 252 # when a single site's verdict looks wrong, read the response it actually got 253 sherlock user123 --site Steam --dump-response 254 255 # run against an unmerged upstream site list, or a local one you have edited 256 sherlock user123 --json https://example.org/data.json 257 sherlock user123 --local 258 ``` 259 260 Interpreting the output: a green `[+]` is a URL that matched the site's detection rule, nothing 261 more. Sites excluded upstream for being unreliable are skipped silently unless you pass 262 `--ignore-exclusions`, which, as its own help text says, brings back false positives along with 263 the coverage. 264 265 How it misleads you: 327 of the 481 entries decide on the HTTP status code alone, so every 266 soft-404 site in the list is a coin toss. Ninety-five entries carry a `regexCheck` that rejects 267 handles which cannot exist on that platform — useful, but it means a handle containing a dot or a 268 dash will be reported as absent from sites that simply do not permit the character, which is a 269 different statement from "nobody registered it". And the `--tor` flag that circulates in older 270 tutorials no longer exists; current releases have `--proxy` only, and passing `--tor` will stop 271 the run with an argparse error rather than quietly running unproxied. 272 273 ### Maigret 274 275 Broader than Sherlock and considerably more useful after the first pass, because it extracts 276 profile fields rather than only reporting existence. It also searches by identifiers other than 277 usernames, which is how you re-enter the search from a numeric ID. 278 279 ```bash 280 pipx install maigret # or: pip install maigret, sudo snap install maigret 281 docker pull soxoj/maigret 282 ``` 283 284 ```bash 285 # default run: the 500 highest-traffic sites 286 maigret user123 287 288 # everything in the database, which is where the long-tail regional sites live 289 maigret user123 -a 290 291 # trade coverage for speed on a first sweep 292 maigret user123 --top-sites 100 293 294 # scope by site tag; run --stats first to see which tags the database actually uses 295 maigret user123 --stats 296 maigret user123 --tags photo,dating 297 maigret user123 --exclude-tags "xx NSFW xx" 298 299 # reports: -H html, -P pdf, -C csv, -T txt, -M markdown, -G graph, -X xmind 300 maigret user123 -H -C --folderoutput ./runs/kestrel 301 302 # JSON for a pipeline; TYPE is 'simple' or 'ndjson' 303 maigret user123 -J ndjson 304 305 # pivot from a profile page: parse it, extract IDs and usernames, search on those 306 maigret --parse https://www.deviantart.com/muse1908 307 308 # search by a stable identifier instead of a handle 309 maigret 112233445566778899001 --id-type gaia_id 310 maigret 12345678 --id-type vk_id 311 312 # combine two name fragments into candidate handles 313 maigret john smith --permute 314 315 # highlight sites whose page body also mentions a term you care about 316 maigret user123 --keywords photography lisbon 317 318 # throttle and proxy: 6,000-odd database entries at default concurrency is a very loud burst 319 maigret user123 -n 20 --timeout 30 --proxy socks5://127.0.0.1:1080 320 maigret user123 --tor-proxy socks5://127.0.0.1:9050 321 ``` 322 323 The valid `--id-type` values are fixed by the source and worth knowing, because they are the exit 324 doors from handle-based reasoning: `username`, `yandex_public_id`, `gaia_id`, `vk_id`, `ok_id`, 325 `wikimapia_uid`, `steam_id`, `uidme_uguid`, `yelp_userid`, `orcid`, `qq_id`, `bilibili_id`. 326 327 Before you trust a long run, measure the list: 328 329 ```bash 330 # tests each site against the known-claimed and known-unclaimed handles in the database 331 maigret --self-check --diagnose 332 333 # and disable the ones that fail, so the next run is quieter 334 maigret --self-check --auto-disable 335 ``` 336 337 How it misleads you: recursive search is on by default, so Maigret will take a username it found 338 on one profile page and start searching for *that*, which is excellent until the profile belongs 339 to someone else and the report quietly becomes a report about two people. Pass `--no-recursion` 340 when you want a clean single-handle run, and read the report's own account-tree rather than the 341 flat list. The `--ai` summary is generated text about your results, not a finding, and should 342 never appear in a write-up. Note also that `--print-not-found` is off by default, so a short 343 report can mean a short list of sites rather than a short list of accounts. 344 345 ### WhatsMyName and Naminter 346 347 [WhatsMyName](https://whatsmyname.app/) is the community-maintained detection dataset — a single 348 `wmn-data.json` of 717 site entries under CC BY-SA 4.0 — rather than a checker. The project 349 removed its bundled scripts in 2023 and now maintains only the data, which several other tools 350 consume. Its detection strings tend to be the most current of any list, because fixing one is a 351 one-line pull request. 352 353 Three inclusion rules shape what it can ever find, and they are the reason an absence here means 354 very little: a site must be publicly accessible without a login, the username must appear in the 355 profile URL, and the site must not swap the username for a numeric ID. Every platform that 356 identifies people by number is therefore structurally invisible to this entire class of tool. 357 358 [Naminter](https://github.com/3xp0rt/Naminter) is the checker built specifically for that dataset, 359 and it is the one to use when you care about the difference between a hit and a maybe. 360 361 ```bash 362 pip install naminter # or: uv tool install naminter 363 docker run --rm -it ghcr.io/3xp0rt/naminter --username john_doe 364 naminter --version 365 ``` 366 367 ```bash 368 # strict detection is the default: status AND string must both match 369 naminter -u user123 370 371 # show every status, not just the hits -- this is the run you actually read 372 naminter -u user123 --filter-all 373 374 # the partials are the pile that needs eyes; the conflicts are a broken site entry 375 naminter -u user123 --filter-partial --filter-conflicting 376 377 # scope by category; the dataset's 20 values include social, coding, gaming, finance 378 naminter -u user123 --include-categories social --include-categories coding 379 naminter -u user123 --exclude-categories "xx NSFW xx" 380 381 # keep the evidence: raw response bodies plus a report, before anything is deleted 382 naminter -u user123 --filter-all --save-response --response-dir ./evidence \ 383 --csv --csv-path hits.csv --html --html-path report.html 384 385 # slow it down and proxy it 386 naminter -u user123 --proxy socks5://127.0.0.1:9050 --timeout 60 --max-tasks 20 387 388 # pin the dataset to a local copy so a run is reproducible months later 389 naminter -u user123 --local-data ./wmn-data.json --local-schema ./wmn-data-schema.json 390 391 # measure the list before you trust it: checks every site against its known accounts 392 naminter --test 393 naminter --test -s GitHub -s Reddit -vvv 394 ``` 395 396 Read the status symbols rather than counting lines. `+` exists, `-` missing, `~+` and `~-` are 397 partial matches where only one of the two criteria fired, `*` is conflicting — the exists and 398 missing indicators both matched, which means the site entry is stale — `?` unknown, `X` a site 399 marked invalid in the data, and `!` a request error. A partial is not a weak hit; it is a site 400 where the tool cannot tell, and the fix is to open the URL. 401 402 How it misleads you: `--impersonate` defaults to `chrome`, so Naminter is deliberately not 403 truthful about what it is. That is usually what you want, but it means a site serving a different 404 page to automation may give a verdict you cannot reproduce with `curl`, and `--impersonate none` 405 will sometimes flip a result. `--max-tasks` defaults to 50, which against 717 sites is a visible 406 burst from one address. And `--filter-exists` is the default, so a quiet run is not a negative 407 result until you have re-run it with `--filter-all`. 408 409 ### Blackbird 410 411 Searches by username and by email from the same binary, which is its reason to exist: a single 412 pass over both identifier types, with permutation built in. Useful as a cross-check when Sherlock 413 and Naminter disagree, because it is a third implementation reading a different list. 414 415 ```bash 416 git clone https://github.com/p1ngul1n0/blackbird 417 cd blackbird 418 pip install -r requirements.txt 419 ``` 420 421 ```bash 422 python blackbird.py --username user123 423 python blackbird.py --email target@example.com 424 425 # several at once, or from a file 426 python blackbird.py -u alice bob charlie 427 python blackbird.py -uf handles.txt 428 python blackbird.py -ef addresses.txt 429 430 # permutation: --permute ignores single elements, --permuteall combines everything 431 python blackbird.py -u john smith --permute 432 python blackbird.py -u john michael smith --permuteall 433 434 # scope by a list property 435 python blackbird.py -u user123 --filter "cat=social" 436 python blackbird.py -u user123 --no-nsfw 437 438 # keep the pages, not just the verdicts 439 python blackbird.py -u user123 --dump --csv --json 440 441 # throttle, proxy, and freeze the site list for a reproducible run 442 python blackbird.py -u user123 --proxy socks5://127.0.0.1:9050 \ 443 --timeout 60 --max-concurrent-requests 10 --no-update 444 ``` 445 446 How it misleads you: it updates its site lists on every run unless you pass `--no-update`, so two 447 runs a week apart are not comparable and neither is reproducible — pin it with `--no-update` the 448 moment a run matters. The `--ai` profiling requires `--setup-ai` and an API key, and it produces 449 characterisation of a person from thin evidence, which is exactly the thing an investigation 450 should not be doing with a language model. Treat the account list as the output and ignore the 451 rest. 452 453 ### socid_extractor 454 455 This is the step that turns a list of URLs into evidence. It parses a profile page or API response 456 and returns a flat dictionary of fields, including the stable internal identifiers a platform uses 457 behind the handle — GAIA ID for Google, Facebook UID, Yandex public ID, Instagram `pk`. Those 458 survive renames, which handles do not. It is the extraction engine Maigret uses, and it is worth 459 running directly on each confirmed hit. 460 461 ```bash 462 pipx install socid-extractor # or: pip install socid-extractor 463 ``` 464 465 ```bash 466 # one profile 467 socid_extractor --url https://www.deviantart.com/muse1908 468 469 # structured, for the case file 470 socid_extractor --url https://github.com/torvalds --json 471 472 # a page you already saved, which is the right way to work from archived evidence 473 socid_extractor --file ./evidence/profile.html --json 474 475 # sites that need a session 476 socid_extractor --url https://example.com/profile --cookies 'sessionid=...; csrftoken=...' 477 socid_extractor --cookie-jar ./cookies.txt --url https://example.com/profile 478 479 # batch runs: skip the HTTP request when the URL matches no known scheme 480 socid_extractor --url https://example.com/foo --skip-fetch-if-no-url-hint 481 ``` 482 483 Output is `key: value` lines, or JSON with `--json`. The fields worth extracting first are any 484 `*_id`, `created_at` and the avatar URL; the rest is context. 485 486 How it misleads you: it reads what the page serves, so a page behind Cloudflare, a login wall or 487 an A/B test returns nothing and that is not the same as the account not existing. The field 488 ontology is normalised across platforms, which means `created_at` from two sites may have been 489 derived two different ways — one from an API field, one from a rendered string — and comparing 490 them to the minute is over-reading. The `--ai-fallback` extractor guesses at pages with no 491 matching scheme; anything it produces is unverified and should be marked as such. 492 493 ### GHunt 494 495 The Google-specific case, included here because a GAIA ID is the most portable stable identifier 496 you are likely to recover, and because `maigret --id-type gaia_id` takes one directly. 497 498 ```bash 499 pipx install ghunt 500 ghunt login # authenticates GHunt to Google; required before the modules work 501 ``` 502 503 ```bash 504 ghunt email target@gmail.com --json user_data.json 505 ghunt gaia 112233445566778899001 --json gaia.json 506 ``` 507 508 The modules are `login`, `email`, `gaia`, `drive`, `geolocate` and `spiderdal`, and `--json` works 509 with all but `login`. Email-side work — validation, breach and registration checks, recovery-hint 510 pivots — belongs on [Email & Phone Number OSINT](/sheets/osint/email-and-phone) rather than here. 511 512 How it misleads you: it runs as an authenticated Google session, so the account you log in with is 513 visible to Google and the results depend on what that account is permitted to see. Use a research 514 account, never a personal one. 515 516 ### Name and handle variant generation 517 518 Enumeration tests the strings you give it, so the strings are the whole game. Two web tools 519 generate them, and both are worth a pass before you conclude a handle is absent. 520 521 [Bellingcat's Name Variant Search](https://bellingcat.github.io/name-variant-search/) generates 522 plausible alternative forms of a person's name — the problem it solves is transliteration and 523 spelling across alphabets, where a registry search fails not because the person is absent but 524 because the name was romanised differently. [NAMINT](https://seintpl.github.io/NAMINT/) works the 525 other way round, taking first, middle and last names and producing username and email 526 combinations. 527 528 Then there are the mechanical variants, which you can generate without a tool: 529 530 ```text 531 Given "John Smith", born 1990, known handle jsmith: 532 separators jsmith j_smith j-smith j.smith 533 order smithj smith_j johns sjohn 534 year suffixes jsmith90 jsmith1990 jsmith09 535 vowel drops jsmth jnsmth 536 leet j5mith jsm1th 537 platform caps usernames over the length limit get truncated differently per site 538 ``` 539 540 Sherlock's `{?}` covers only the three separators. Maigret's `--permute` and Blackbird's 541 `--permute` / `--permuteall` combine name fragments you supply. Nothing generates the year 542 suffixes or the vowel drops for you, so build that list by hand and feed it as multiple 543 positional usernames. 544 545 How this misleads you: variant generation multiplies your request volume by the number of 546 variants, and it multiplies your false-positive count by the same factor. Ten variants against 717 547 sites is 7,170 requests and roughly ten times the soft-404 noise. Generate variants, but run them 548 scoped to a category or a site list rather than across everything. 549 550 ## Enumerators at a glance 551 552 | Tool | Site list | Detection rule | Extracts profile data | Reports | 553 | --- | --- | --- | --- | --- | 554 | [Sherlock](https://github.com/sherlock-project/sherlock) | Own `data.json`, 481 sites | Status code on 327 of 481; message or response URL on the rest | No | csv, xlsx, txt | 555 | [Maigret](https://github.com/soxoj/maigret) | Own database, top 500 by default, `-a` for all | Per-site, plus recursive search on extracted usernames | Yes, via socid_extractor | html, pdf, csv, txt, md, graph, xmind, json, neo4j | 556 | [Naminter](https://github.com/3xp0rt/Naminter) | WhatsMyName, 717 sites | Strict `all` by default; exposes partial and conflicting states | No | csv, json, html, pdf, raw responses | 557 | [Blackbird](https://github.com/p1ngul1n0/blackbird) | Own list, auto-updating | Per-site | Dumps HTML with `--dump` | csv, json, pdf | 558 | [WhatsMyName web](https://whatsmyname.app/) | WhatsMyName, 717 sites | As the dataset defines | No | csv | 559 | [socid_extractor](https://github.com/soxoj/socid-extractor) | Not an enumerator; 250+ parse schemes | n/a | Yes, including stable IDs | json | 560 561 Pick on the second and third columns, not the first. Sherlock for a fast look, Naminter when the 562 count has to be defensible, Maigret when you want the profile fields in the same pass, Blackbird 563 as the third opinion, socid_extractor on every hit that survives. 564 565 ## Tool reference 566 567 | Tool | What it does | Cost | 568 | --- | --- | --- | 569 | [192](http://www.192.com/) | Searching for someone's address in the UK, phone number and who they live with according to electoral rolls. | free | 570 | [Bellingcat Name Variant Search](https://bellingcat.github.io/name-variant-search/) | Quickly search many variants of a person's name on Google | free | 571 | [Blackbird](https://github.com/p1ngul1n0/blackbird) | Check usernames and email addresses on websites and social networks | free | 572 | [Epieos](https://tools.epieos.com/holehe.php) | Checks where an email has been used. Based on Holehe. | paid | 573 | [Ghunt](https://github.com/mxrch/GHunt) | A command line tool for obtaining information about Google accounts. | free | 574 | [Maigret](https://github.com/soxoj/maigret) | Maigret is a Python script that retrieves user information by searching for usernames across various websites and social media platforms. | free | 575 | [NeutrOSINT](https://github.com/Kr0wZ/NeutrOSINT) | A tool for investigating Proton Mail addresses. | free | 576 | [Sherlock](https://github.com/sherlock-project/sherlock) | Allows a user to search for the presence of specific usernames across more than 400 websites and social networks. | free | 577 | [WhatsMyName](https://whatsmyname.app/) | Search for usernames on several hundred platforms | free | 578 579 ## Pitfalls 580 581 - **A 200 is not a profile.** 178 of WhatsMyName's 717 entries record HTTP 200 for an account that 582 does not exist, and 327 of Sherlock's 481 entries decide on the status code alone. On those 583 sites a hit is a coin toss until you open the URL. 584 - **`--mode any` is not "more thorough".** A loose detection rule can promote up to a quarter of 585 the list to EXISTS without a single real account being involved. The extra hits are not weak 586 evidence; they are not evidence. 587 - **Partials are the interesting pile, not the discard pile.** A PARTIAL_EXISTS means the status 588 matched and the string did not, or the reverse, which usually means the site redesigned. Those 589 are the sites where a real account is most likely to be hiding from the tool. 590 - **A CONFLICTING result is a bug report, not a finding.** Both the exists and missing indicators 591 matched, so the site entry is stale. Fix it upstream or ignore the site; do not average the two. 592 - **The dataset cannot see numeric-ID platforms at all.** WhatsMyName only includes sites where 593 the username appears in the URL and is not transformed into an internal ID. Absence from every 594 enumerator you own is therefore consistent with a large presence on platforms that identify 595 people by number. 596 - **Handles are recycled.** A handle released by one account and claimed by another means the 597 2019 archive and the live profile are different people wearing the same name. Check creation 598 dates before you treat an archived page and a current one as the same account. 599 - **Display name and handle are not independent evidence.** Both derive from the real name. 600 Multiplying them together is how a 1:2 prior becomes a confident, wrong answer. 601 - **A matching avatar is only as good as the image's rarity.** An identical bitmap that also 602 appears on 4,000 other profiles is worth nothing. Reverse-search the avatar before you score 603 it — see [Reverse Image Search](/sheets/osint/reverse-image-search). 604 - **Self-check the list before you trust a long run.** `maigret --self-check --diagnose` and 605 `naminter --test` both measure the detection rules against known accounts. A list you have not 606 measured gives you a hit count with no error bar. 607 - **Browser impersonation changes the answer.** Naminter impersonates Chrome by default and 608 Blackbird can be made to; a verdict you cannot reproduce with a plain client is a verdict that 609 depends on a header. Record which mode produced it. 610 - **Enumerators are loud.** 717 sites at Naminter's default 50 concurrent requests is a burst from 611 one address in a few seconds. Proxy it, lower `--max-tasks` or `-n`, and assume the target can 612 see the pattern if they operate any of the sites. 613 - **Variant sweeps multiply the noise, not just the coverage.** Ten variants is ten times the 614 requests and ten times the soft-404 hits. Scope them by category. 615 - **Coverage is skewed.** These lists are heavy on English-language and global platforms and thin 616 on regional ones. Absence from a tool's results is not absence from the internet. 617 - **Paid aggregators recycle stale data.** A "verified" result from a people-search service is 618 often a years-old scrape. Check when the underlying data was collected — see 619 [People Search & Public Records](/sheets/osint/people-search). 620 - **The AI summary is not a finding.** Maigret's `--ai` and Blackbird's `--ai` produce generated 621 prose about your results. It cites nothing, and it will characterise a person from three data 622 points. Keep it out of the write-up. 623 - **Archive before you cite.** Profiles are deleted while you are looking at them. Capture with 624 `--save-response` or `--dump` and preserve properly — see 625 [Archiving & Evidence Preservation](/sheets/osint/archiving-and-evidence). 626 627 ## Worked example 628 629 One datum: the handle `kestrel_ait`, seen once in a forum signature. No name, no email, no claim 630 about who it belongs to. The question asked is narrow — does the person behind the GitHub account 631 with that handle also hold the Chess.com account with that handle. 632 633 1. **Sweep the separators.** Sherlock's `{?}` expands to `_`, `-` and `.`, which covers the three 634 mechanical variants in one invocation: 635 636 ```bash 637 sherlock 'kestrel{?}ait' --csv --folderoutput ./runs/kestrel 638 ``` 639 640 `kestrel-ait` and `kestrel.ait` return two hits each, both on soft-404 sites, both empty pages 641 when opened. The underscore form is the live one and the rest of the work uses it. 642 643 2. **Run the strict pass and read every status.** Naminter against the full WhatsMyName dataset, 644 with responses preserved: 645 646 ```bash 647 naminter -u kestrel_ait --filter-all --mode all \ 648 --save-response --response-dir ./evidence \ 649 --csv --csv-path kestrel.csv --max-tasks 20 650 ``` 651 652 ```text 653 [+] GitHub https://github.com/kestrel_ait 654 [+] Chess.com https://api.chess.com/pub/player/kestrel_ait 655 [+] Last.fm https://www.last.fm/user/kestrel_ait 656 [+] Mastodon (infosec.exchange) 657 [+] SoundCloud https://soundcloud.com/kestrel_ait 658 [+] Patreon https://www.patreon.com/kestrel_ait 659 [+] Gravatar https://en.gravatar.com/kestrel_ait.json 660 [+] Steam https://steamcommunity.com/id/kestrel_ait 661 [+] Keybase https://keybase.io/_/api/1.0/user/lookup.json?usernames=kestrel_ait 662 663 exists 9 partial 23 conflicting 4 error 6 missing 675 total 717 664 ``` 665 666 3. **Quantify the ambiguity rather than guessing at it.** The same run with the loose rule: 667 668 ```bash 669 naminter -u kestrel_ait --filter-exists --mode any --max-tasks 20 670 ``` 671 672 returns **36** hits. The 27 extra are exactly the 23 partials plus the 4 conflicts, promoted by 673 a rule that accepts a status match on its own. Nothing was discovered between the two runs. If 674 the report had been written from the `any` run it would have claimed 36 accounts, of which at 675 most 9 were ever candidates. 676 677 4. **Open all nine.** Steam and Keybase are the soft-404 sites from the arithmetic above: both 678 return 200 and both pages are empty. Gravatar resolves to a default identicon with no profile. 679 Patreon is a reserved-but-unused page. That leaves five: GitHub, Chess.com, Last.fm, SoundCloud 680 and the Mastodon account. 681 682 5. **Estimate `k`.** Last.fm has scrobbles since 2014, all of one genre, with a listed location in 683 another country. SoundCloud is a label account with a logo avatar and a contact address on a 684 company domain. Both are positively excluded. So at least three distinct humans or entities 685 hold this handle: the GitHub owner, the Last.fm owner, the SoundCloud owner. 686 687 ```text 688 k = 3 -> prior that Chess.com is the GitHub person = 1/3 = 0.33 (odds 1:2 against) 689 ``` 690 691 6. **Pull the stable identifiers.** Both remaining platforms publish an API: 692 693 ```bash 694 socid_extractor --url https://github.com/kestrel_ait --json 695 curl -s https://api.chess.com/pub/player/kestrel_ait 696 ``` 697 698 ```json 699 { 700 "player_id": 196341287, 701 "@id": "https://api.chess.com/pub/player/kestrel_ait", 702 "username": "kestrel_ait", 703 "status": "basic", 704 "avatar": "https://images.chesscomfiles.com/uploads/v1/user/196341287.9f1c2a7e.200x200o.png", 705 "joined": 1696153260, 706 "last_online": 1759512000, 707 "followers": 4, 708 "country": "https://api.chess.com/pub/country/GB" 709 } 710 ``` 711 712 GitHub returns `"id": 146902331` and `"created_at": "2023-10-01T10:12:04Z"`. Neither ID is the 713 other, and no identifier is shared, so there is no account-level link. This stays an inference. 714 715 7. **Do the timestamp arithmetic properly.** `joined` is a Unix epoch and has to be converted 716 before it means anything: 717 718 ```bash 719 python3 -c "import datetime as dt; print(dt.datetime.fromtimestamp(1696153260, dt.timezone.utc))" 720 # 2023-10-01 09:41:00+00:00 721 ``` 722 723 ```text 724 Chess.com joined 2023-10-01 09:41:00 UTC 725 GitHub created_at 2023-10-01 10:12:04 UTC 726 gap 31 minutes 727 ``` 728 729 8. **Score the avatar, and check its rarity first.** Both avatars are a cropped photograph of a 730 bird of prey. Reverse-image searching it returns no other instance, so it is not a stock image: 731 732 ```python 733 from PIL import Image 734 import imagehash 735 736 gh = imagehash.phash(Image.open("evidence/github_avatar.png")) 737 cc = imagehash.phash(Image.open("evidence/chesscom_avatar.png")) 738 print(gh, cc, gh - cc) # c3a1e4b2d8f09156 c3a1e4b2d8f09156 0 739 ``` 740 741 A 64-bit pHash distance of 0 on a non-stock image is the strongest single piece of evidence in 742 the file. 743 744 9. **Multiply, with the numbers written down.** 745 746 ```text 747 prior odds 1 : 2 748 avatar, distance 0, image found nowhere else x 60 749 registration 31 minutes apart x 20 750 --------- 751 posterior odds 1200 : 2 = 600 : 1 p = 0.998 752 ``` 753 754 Those two likelihood ratios are judgements, not measurements, and they are stated so that 755 somebody can halve them. Even halved twice the answer does not change, which is the useful 756 property of writing them down. The display name was deliberately not counted: both accounts 757 read "K. Aitken", and that is the same evidence as the handle, not a second piece. 758 759 What you can assert: two accounts, GitHub `id` 146902331 and Chess.com `player_id` 196341287, 760 registered 31 minutes apart on 1 October 2023 and carrying a byte-identical avatar that appears 761 nowhere else, are the same person with posterior odds of roughly 600:1 given a prior of 1:2 from 762 three observed holders of the handle. No shared platform identifier was recovered, so this is an 763 inference from independent evidence and not an account-level link, and the write-up should say so 764 in those words. 765 766 What would falsify it: a fourth and fifth holder of the handle, which would lower the prior; any 767 instance of that avatar image elsewhere on the internet, which would collapse its likelihood ratio 768 from 60 to about 1 and take the conclusion with it; or a Chess.com `joined` timestamp read in the 769 wrong zone — the epoch is UTC, and treating it as local time would move the gap from 31 minutes to 770 an hour or more and remove the clustering argument entirely. The measurement that carries the most 771 risk is the avatar's rarity, not its hash distance. The distance is exact and tolerates nothing 772 above about 2 before it stops meaning "same file"; the rarity is a search result that can be wrong 773 in one direction only, and it is the one claim here that must be re-checked before publication. 774 775 ## Broader catalogues 776 777 - [Username OSINT](https://tools.osintnewsletter.com/tool-categories/username-osint) 778 - [People OSINT](https://tools.osintnewsletter.com/tool-categories/people-osint) 779 780 781 ## More tools 782 783 Further tools for this area from the OSINT Newsletter Tools Library ([Username OSINT](https://tools.osintnewsletter.com/tool-categories/username-osint)), excluding those already listed above. 784 785 | Tool | What it does | 786 | --- | --- | 787 | [Analyst Research Tools](https://analystresearchtools.com/) | A browser-based investigative research platform pulling together a collection of free lookup tools to help you gather, organise… | 788 | [DeepFind.Me Username Search](https://deepfind.me/tools/social-media/username-search) | Username search tool that scans multiple online platforms to identify where a specific username appears across social media… | 789 | [Hippie OSINT Toolkit](https://osint.hippie.cat/) | An OSINT web toolkit that allows you to reverse search a domain, a TikTok post, an image, or username (and more). | 790 | [LoLArchiver](https://lolarchiver.com/) | A niche OSINT tool designed to archive and surface user-related data across gaming and online platforms like League of Legends… | 791 | [NAMINT](https://seintpl.github.io/NAMINT/) | A web app that generates multiple username, email, and naming combinations from a person’s first, middle, and last name. | 792 | [OSINT Industries](https://app.osint.industries/) | An all-encompassing OSINT platform that gathers and correlates publicly available digital data such as emails, domains, phone… | 793 | [Osintly](https://osint.ly/) | A unified OSINT platform that searches for information across pseudonyms, email addresses, phone numbers, IP addresses, domains… | 794 | [PSNProfiles](https://psnprofiles.com/) | A PlayStation trophy tracking and player statistics platform that allows users to view public PlayStation Network profiles… | 795 | [Revealer.US username lookup](https://revealer.us/) | An exposure-intelligence and OSINT tool that lets you search for information connected to an email address, username, or phone… | 796 | [Sherlock Eye](https://www.sherlockeye.io/) | AI-powered OSINT assistant that automates investigations, links data, and surfaces insights fast. | 797 | [TheBigBrother](https://github.com/chadi0x/TheBigBrother) | All-in-one OSINT toolkit that helps track usernames across the web, investigate emails and domains, extract metadata, and follow… | 798 | [Threat Actor Username Search](https://threatactorusernames.com/) | A simple OSINT tool that checks whether a username has been observed on known cybercriminal forums and underground platforms. | 799 800 ## Sources 801 802 Both catalogues below are maintained by other people and are considerably larger than 803 this page. Use them as the canonical index; this sheet is a working route through them. 804 805 - [Bellingcat's Online Investigation Toolkit](https://bellingcat.gitbook.io/toolkit) — ~340 tools, each with its own 806 review page covering cost, difficulty, requirements and limitations. 807 - [OSINT Newsletter Tools Library](https://tools.osintnewsletter.com) — ~280 tools, organised by investigative goal. 808 809 Neither publishes a licence, so nothing here is copied from them: tool names, one-line 810 descriptions, cost flags and links are catalogue facts, and the method and commentary are 811 this site's own. See [credits](/credits).