commit 6d7bafcfec145c6a61ff6a9e2c35b0f6c5db3d3c
parent 18e9f45dce9a9126f5aab0732fb4187f1eec5dfe
Author: DAEMON <zer0sec.xp@icloud.com>
Date: Sun, 4 Oct 2026 07:13:37 +0100
Expand verified OSINT guides and add the global signal field
Diffstat:
16 files changed, 3860 insertions(+), 220 deletions(-)
diff --git a/src/content/sheets/osint/archiving-and-evidence.md b/src/content/sheets/osint/archiving-and-evidence.md
@@ -701,13 +701,18 @@ Then the same URL by hand into [archive.today](https://archive.today), and its s
capture time into `PROVENANCE.txt`. Submitting last is deliberate: both submissions are public, so
if the account notices and deletes the post, you already hold the capture.
-What this run establishes: the page as it stood at a timestamped moment, the clip with the
+What you can assert: the page as it stood at a timestamped moment, the clip with the
platform's own metadata, and a change history from Wayback showing the post was edited after the
video was uploaded. What it does not establish: that the footage shows what the caption claims, or
where or when it was filmed. Those are
[Image & Video Forensics](/sheets/osint/image-video-forensics) and
[Geolocation](/sheets/osint/geolocation) questions, and no amount of archiving substitutes for them.
+What would falsify it: any manifest entry that fails verification, a timestamp token whose digest
+does not match the manifest, or a WARC/WACZ log showing that the claimed resource failed and only
+an error page was captured. The public archive copy is corroboration; the local capture, hashes
+and independently verifiable timestamps carry the claim.
+
## Broader catalogues
- [Archiving OSINT](https://tools.osintnewsletter.com/tool-categories/archiving-osint)
diff --git a/src/content/sheets/osint/companies-and-finance.md b/src/content/sheets/osint/companies-and-finance.md
@@ -584,7 +584,7 @@ curl -s -A "$UA" -G 'https://efts.sec.gov/LATEST/search-index' \
# one 10-K exhibit, 2023, listing it as a counterparty of a listed miner
```
-The defensible chain: invoice name → UK company 09876543 (Companies House, retrieved with date) →
+What you can assert: invoice name → UK company 09876543 (Companies House, retrieved with date) →
corporate PSC in BVI (UK filing, self-declared) → named individual with 100% ownership from 2014
(Aleph, ICIJ collection, snapshot date stated) → confirmed independently in the ICIJ public
database via the same intermediary → named as a counterparty in a 2023 SEC exhibit. Six sources,
@@ -596,6 +596,11 @@ rather than the person. "Beneficial ownership is not currently published in the
Islands; the most recent evidenced owner is X as of 2014" is the honest sentence, and it is a
stronger finding than an unqualified assertion.
+What would falsify it: a later primary filing naming a different beneficial owner, the Aleph and
+ICIJ records resolving to different entities once registration numbers are compared, or the SEC
+exhibit referring to a namesake rather than company 09876543. The 2014 ownership edge is the
+oldest and weakest link; date it explicitly and replace it when a newer primary record appears.
+
## Broader catalogues
- [Public Records OSINT](https://tools.osintnewsletter.com/tool-categories/public-records-osint)
diff --git a/src/content/sheets/osint/conflict-and-environment.md b/src/content/sheets/osint/conflict-and-environment.md
@@ -587,7 +587,7 @@ Pull both as B12/B8A/B4 GeoTIFFs through `/api/v1/process`, hash each, and recor
The 14 March image shows a dark scar and a collapsed roof span on one silo where the 9 March image
shows an intact structure.
-What is now defensible: a thermal anomaly at the coordinate on the night of 11 March, from a sensor
+What you can assert: a thermal anomaly at the coordinate on the night of 11 March, from a sensor
and not a report; no comparable detection in that pixel across the preceding year; and structural
change visible between two named Sentinel-2 scenes bracketing the date. Three legs, two of them
independent of all media reporting, each tied to a query, a retrieval time and a hash.
@@ -597,6 +597,11 @@ drone, an accident or deliberate arson, and attribution needs the remnant. If a
debris surfaces, that goes to [OSMP](https://osmp.ngo/) for a markings comparison, and the
identification is stated as consistent-with until a lot number is legible.
+What would falsify it: a second pre-event image showing the roof already damaged, a FIRMS quality
+flag or geolocation error placing the detections outside the terminal, or a longer baseline showing
+routine industrial heat at the same pixel. The imagery pair carries the structural-change claim;
+the thermal detections only narrow the time window and do not identify a cause.
+
## Broader catalogues
- [Conflict OSINT](https://tools.osintnewsletter.com/tool-categories/conflict-osint)
diff --git a/src/content/sheets/osint/data-analysis-and-visualisation.md b/src/content/sheets/osint/data-analysis-and-visualisation.md
@@ -646,13 +646,18 @@ Datawrapper, with the two-day collection gap on 5–7 March drawn as a shaded ba
line reading "12,386 posts collected 2026-02-01 to 2026-03-31; 38 timestamps unparsed; four company
name variants merged".
-What this run establishes: a corrected counterparty ranking, a reproducible cleaning record
+What you can assert: a corrected counterparty ranking, a reproducible cleaning record
including one rejected merge, and a quantified timing relationship between two accounts. What it
does not establish: that `acct_f` and `acct_k` are operated by the same person, or coordinated at
all — posting within a minute is consistent with coordination, with both reacting to the same
trigger, and with a scheduler neither of them controls. The chart cannot distinguish those, and
saying so is the finding's credibility.
+What would falsify it: rerunning the saved cleaning steps changes the row count, the sixty-second
+pairing disappears when collection gaps and duplicate posts are removed, or the relationship is
+fully explained by a shared external event. The timing threshold is the sensitive measurement;
+publish the result at several thresholds rather than treating sixty seconds as a natural law.
+
## Broader catalogues
- [Data Extraction OSINT](https://tools.osintnewsletter.com/tool-categories/data-extraction-osint)
diff --git a/src/content/sheets/osint/email-and-phone.md b/src/content/sheets/osint/email-and-phone.md
@@ -4,9 +4,9 @@ description: "Validate an address or number, find the accounts attached to it, a
category: osint
subcategory: "People & Identity"
tags: [osint, email, phone, pivoting]
-tools: [epieos, ghunt, holehe, phoneinfoga]
+tools: [epieos, ghunt, holehe, zehef, phoneinfoga, phunter, ignorant, libphonenumber, swaks, gravatar]
difficulty: intermediate
-updated: 2026-09-28
+updated: 2026-10-04
references:
- name: "Bellingcat's Online Investigation Toolkit"
url: "https://bellingcat.gitbook.io/toolkit"
@@ -20,6 +20,37 @@ references:
license: none
relation: derived
note: "Second tool catalogue, cross-checked against the above."
+ - name: "libphonenumber FAQ"
+ url: "https://github.com/google/libphonenumber/blob/master/FAQ.md"
+ author: "Google"
+ license: Apache-2.0
+ relation: link-only
+ note: "Upstream semantics for validity, carrier mapping and mobile number portability."
+ - name: "PhoneInfoga documentation"
+ url: "https://sundowndev.github.io/phoneinfoga/"
+ author: "Sundowndev"
+ relation: link-only
+ note: "Install methods, CLI flags and the scanner list were verified against this."
+ - name: "GHunt"
+ url: "https://github.com/mxrch/GHunt"
+ author: "mxrch"
+ relation: link-only
+ note: "Subcommand names and the three login methods were verified against this README."
+ - name: "Gravatar developer documentation: avatar images"
+ url: "https://docs.gravatar.com/api/avatars/images/"
+ author: "Automattic"
+ relation: link-only
+ note: "The d= and s= query parameters used for the existence probe."
+ - name: "Gravatar developer documentation: creating the hash"
+ url: "https://docs.gravatar.com/api/avatars/hash/"
+ author: "Automattic"
+ relation: link-only
+ note: "The trim, lower-case and SHA-256 recipe, and the worked example digest reproduced here."
+ - name: "Ofcom: telephone numbers for drama"
+ url: "https://www.ofcom.org.uk/phones-and-broadband/phone-numbers/numbers-for-drama/"
+ author: "Ofcom"
+ relation: link-only
+ note: "The UK number ranges reserved for fiction, used in the worked example."
---
## What this covers
@@ -28,80 +59,556 @@ An email address or phone number is usually the strongest pivot in an investigat
platforms treat them as account keys. From one address you can often reach a display name, a
profile photo, a set of registered services and sometimes a physical location.
+What makes identifier work different from name work is that the answers are mostly binary and
+mostly cheap. A name returns a ranked list of maybes; an identifier returns yes or no per service.
+That is also the trap: a "no" from a service is four different claims wearing the same word, and
+most of the mistakes on this page come from recording the wrong one.
+
## Method
-1. **Validate before enriching.** Confirm the address or number is real and routable. Time spent
- enriching a typo is wasted.
-2. **Check what it is registered to.** Account-existence checks against signup and
- password-reset flows reveal which services know the identifier.
-3. **Pull provider-side metadata.** Google accounts in particular leak a display name, a profile
- photo, and public calendar and review activity.
-4. **Pivot outward.** Search the identifier as a plain string — in code search, paste sites,
- breach indexes and the target's own site. Addresses get committed to repositories constantly.
-5. **Convert to a name, then switch approach.** Once you have a name, you are doing
+1. **Validate before enriching.** Confirm the address or number is structurally real and routable.
+ Time spent enriching a typo, a placeholder or a drama-range number is wasted, and the check
+ costs nothing.
+2. **Normalise, then keep the original.** Decide which spellings are the same mailbox before you
+ count hits or deduplicate a list. Record the string as found as well as the normalised form —
+ the as-found spelling is what you search for in dumps and code.
+3. **Take the free, non-interactive signals first.** DNS records, a Gravatar hash, provider-side
+ profile data. None of this touches the owner and none of it can be rate-limited into a false
+ negative.
+4. **Then the registration oracles, knowing the cost.** Account-existence checks against signup
+ and password-reset flows reveal which services know the identifier, and some of them generate
+ mail or SMS to the target.
+5. **Pivot outward as a plain string.** Code search, paste sites, breach indexes, the target's own
+ site, PDF metadata. Addresses get committed to repositories constantly, and numbers get left in
+ vCards and footers.
+6. **Pivot inward to durable keys.** An address can change; a Google Gaia ID, a Telegram user ID
+ or an internal account ID cannot. Trade the identifier you were given for the one the platform
+ uses internally.
+7. **Convert to a name, then switch approach.** Once you have a name, you are doing
[people search](/sheets/osint/people-search), not identifier work.
+The judgement calls: decide before you start whether this is a passive engagement, because that one
+decision rules out half the tools below and no tool will remind you. Decide whether you need the
+subscriber or the account — a recovery phone, a work number and a partner's handset all answer
+"what number is attached to this account" and none of them answers "what number does this person
+carry". And decide what a negative is worth to you, because the cheapest way to be wrong here is to
+write down "not registered" when what happened was a rate limit.
+
+## The existence oracle
+
+Everything on the email side reduces to asking some service whether it knows a string, and the
+useful part is knowing exactly which string you asked about. Providers normalise differently, so
+the same mailbox has many spellings and every one of them is a distinct key to every tool, hash and
+dataset you will use.
+
+Gmail ignores dots in the local part — Google's own help states that if a sender "added dots to
+your address, you'll still get that email" — and, like any provider implementing subaddressing,
+treats everything from a `+` to the `@` as a tag on the same mailbox. So:
+
+```text
+d.whitfield@example.com \
+dwhitfield@example.com > one mailbox, if the provider is Gmail
+dwhitfield+shop@example.com /
+
+...and three unrelated keys, everywhere else:
+
+sha256("d.whitfield@example.com") = fc9518a0411a19b071630d4b153e4774e0697053e4015e5eb9e1ab5c934021a3
+sha256("dwhitfield@example.com") = 6a2ae976758ac8100467b0bea363301fcb6fe756c8262b95336be1d03dcbd919
+sha256("dwhitfield+shop@example.com") = bbb5413ad03bc64773f79b7804cd4868f34e7fd7a5fa2e1808617ed45289369f
+```
+
+That matters immediately, because the cheapest non-interactive oracle in this whole area is a hash
+lookup. Gravatar's documented recipe is: trim leading and trailing whitespace, force the string to
+lower case, hash it with SHA-256, and request `https://gravatar.com/avatar/<hash>`. The `d=404`
+parameter turns the avatar endpoint into an existence test — ask for a default image of `404` and
+you get an HTTP status instead of a fallback picture.
+
+```bash
+# the recipe, reproduced against the example in Gravatar's own documentation
+printf '%s' "myemailaddress@example.com" | shasum -a 256
+# 84059b07d4be67b806386c0aad8070a23f18836bbaae342275dc0a83414c32ee
+
+# the probe: 200 means an avatar exists for that exact string, 404 means it does not
+H=$(printf '%s' "d.whitfield@example.com" | shasum -a 256 | cut -d' ' -f1)
+curl -s -o /dev/null -w '%{http_code}\n' "https://gravatar.com/avatar/$H?d=404&s=200"
+# 404
+```
+
+Nothing reaches the owner, there is no rate limit worth worrying about, and a 200 hands you a
+photograph plus, often, a WordPress-ecosystem account. Older code and older datasets used MD5 of
+the same lower-cased string, which is why historic breach corpora and forum databases are full of
+MD5 avatar URLs; the current documentation specifies SHA-256, so compute both when you are matching
+against old data. A 404 speaks only for the exact string you hashed — run it again for the
+dot-stripped and tag-stripped spellings before you call it absent.
+
+The interactive oracles are the signup and password-reset flows, and their answers fall into four
+classes that tools flatten into two:
+
+```text
+what the flow says what you may record
+-------------------------------------------- ------------------------------------------
+"we have sent a reset link to that address" the account exists, on this spelling
+"no account found for that address" no account, on this spelling only
+the same reassuring page for every input nothing -- the site is enumeration-hardened
+429, captcha, timeout, silence nothing -- and you are now rate-limited
+```
+
+`holehe` and `ignorant` both print `[+]` for the first, `[-]` for the second and `[x]` for the
+fourth, and the third is indistinguishable from the second unless you test with a control address
+you know does not exist. Always run one. A `[-]` column with no control is an unmeasured
+instrument.
+
+## E.164 arithmetic
+
+The phone side has less to probe and more to compute. E.164 is the international format every tool
+expects, and it decomposes into exactly three parts with a hard length cap:
+
+```text
++44 7700 900123
+ | | |
+ | | +-- subscriber number
+ | +-------- national destination code (here, a UK mobile block)
+ +------------ country code (1-3 digits)
+
+E.164 allows at most 15 digits after the plus:
+ 44 + 7700900123 = 2 + 10 = 12 digits, inside the limit
+```
+
+Two separate questions get confused here, and libphonenumber keeps them apart deliberately.
+*Possible* means the digit count is plausible for that country. *Valid* means the number matches a
+block the numbering plan actually defines. Neither means the number is assigned to a subscriber,
+and nothing short of ringing it or paying for an HLR lookup establishes that it is in service.
+
+```text
++1 415 555 0123 possible: True valid: True <- and reserved for fiction
++44 7700 900123 possible: True valid: False <- Ofcom drama range
++44 7700 9001 possible: True valid: False <- IS_POSSIBLE_LOCAL_ONLY (4)
++44 770090012345 possible: False valid: False <- TOO_LONG (ValidationResult 3)
+```
+
+The first line is the one to remember. `+1 415 555 0123` passes validation because `555-01xx` sits
+inside a legitimate NANP structure, while the North American numbering plan reserves `555-0100`
+through `555-0199` for fictitious use. Ofcom is more honest about its equivalents and publishes
+them: `07700 900000` to `900999` for mobile, `020 7946 0000` to `0999` for London, and similar
+thousand-number blocks for other cities, with the note that "the numbers will not be allocated to
+communications providers in the foreseeable future". Memorise both. A number from either space in
+a dataset tells you the record is synthetic, and it tells you that for free in the first second of
+the investigation.
+
+The third line hides a second trap. `is_possible_number` returns `True` there not because the
+length is right but because the library counts a number that could be dialled locally, without its
+area code, as possible — the docstring says so outright. The reason code is
+`IS_POSSIBLE_LOCAL_ONLY`, which is `4`; `TOO_SHORT` is `2` and `TOO_LONG` is `3`. The boolean
+collapses all of that, so call `is_possible_number_with_reason` when you intend to write down why
+a number failed rather than that it did.
+
+The other arithmetic worth having is what a masked recovery number is actually worth. Reset screens
+show you a tail — "we will text ●●●●●●23" — and the temptation is to enumerate forwards:
+
+```text
+known prefix from a CV footer : 07700 9 (5 of 10 national digits)
+masked tail from a reset page : 23 (2 of 10)
+unknown digits : 10 - 5 - 2 = 3
+candidates : 10^3 = 1000 <- far too many to probe quietly
+```
+
+So go the other way. Take a number you already have from an independent source and test whether it
+*fits* the mask. That is a confirmation with a known false-positive rate: a two-digit tail matches
+by chance one time in a hundred, which is weak on its own and strong in combination. A Google
+two-digit tail and a courier SMS three-digit tail that both fit the same candidate is a one in
+100,000 coincidence, and that is the kind of statement you can defend.
+
## Email
### Epieos
-The first tool to reach for. Given an address it returns Google and Microsoft account data,
-linked services, and gravatar hits, without sending anything to the address.
+The first tool to reach for, because it is one query, it runs server-side and nothing it does
+reaches the address. Enter the address at [epieos.com](https://epieos.com/) and it returns the
+accounts it can see attached to the string — Google and Microsoft account data, a Gravatar hit, and
+a spread of consumer services. The Google portion is the valuable half: a public account ID, the
+profile photo, and whatever Maps or calendar activity the owner left visible.
+
+```text
+1. Paste the address. Do the dot-stripped and tag-stripped spellings separately; the
+ back end is matching strings, not mailboxes.
+2. Read the Google block first and copy the account ID out of it. That is the Gaia ID,
+ and it is the durable key -- feed it to `ghunt gaia` rather than re-querying the
+ address later.
+3. Treat an empty module as "no answer", not "no account". The result page does not
+ distinguish a service that said no from a service that failed or was not checked on
+ your tier.
+4. Archive the result page before you close it. These lookups are a snapshot of someone's
+ privacy settings on one day and are not reproducible later.
+```
-Enter the address at [epieos.com](https://epieos.com/). No notification reaches the owner.
+The free lookup returns a subset of the modules and the rest sits behind a paid tier, which is the
+usual source of a phantom negative: a module you did not pay for looks exactly like a module that
+found nothing. Capture evidence properly — see
+[Archiving & Evidence](/sheets/osint/archiving-and-evidence) — because "Epieos said so last Tuesday"
+is not a citation.
### GHunt
-Deeper on Google specifically: account ID, display name, profile photo history, public Maps
-reviews and photos, YouTube channel, and calendar if public.
+Deeper on Google specifically, and the tool that converts an address into Google's internal account
+ID. Install it with pipx and authenticate once; `ghunt login` offers three routes, and the
+listening-mode handshake with the GHunt Companion browser extension is the one that does not
+involve pasting cookies by hand.
```bash
pipx install ghunt
-ghunt login # needs your own Google session cookies
+ghunt login # listening mode, base64 cookies, or manual entry
+
+# the subcommands, from --help, before you guess at one
+ghunt -h # {login,email,gaia,drive,geolocate,spiderdal}
ghunt email target@gmail.com
-ghunt gaia 1234567890123456789 # pivot on the internal Google account ID
+ghunt email target@gmail.com --json target_email.json
+
+# the internal account ID, which survives an address change
+ghunt gaia 1234567890123456789 --json target_gaia.json
+
+# a Drive file or folder ID: owner, permissions, and the other files it reveals
+ghunt drive 1A2b3C4d5E6f7G8h9I0jKlMnOpQrStUv
+
+# a BSSID, for when the pivot is a wireless network rather than a person
+ghunt geolocate -b 00:11:22:33:44:55
```
-`ghunt login` requires authenticating with an account you control — use a throwaway, because
-this is exactly the kind of automation Google suspends accounts for.
+Read the output in that order: the Gaia ID first, because it is the thing you cite and the thing
+you pivot on; then the profile photo URL, which goes straight into
+[reverse image search](/sheets/osint/reverse-image-search); then the service-by-service blocks. Use
+`--json` on every run you intend to quote. The terminal rendering is for reading and the JSON is
+for keeping, and the two are not equally easy to re-derive in six months.
+
+How it misleads you: every query runs as an authenticated Google session, so use an account you are
+willing to lose, and expect the unusual request volume to be exactly what Google's abuse systems
+are built to notice. More subtly, an empty module reflects the target's visibility settings rather
+than their absence — a person with a locked-down profile and a person with no Google account
+produce output that looks similar at a glance. And the Drive module reports what your session can
+see, so a file shared with your throwaway account and a genuinely public file are not distinguished
+for you.
### holehe
-Checks an address against 100+ sites by watching how their password-reset and registration flows
-respond.
+The broad sweep: it checks an address against 120-plus sites by watching how their registration and
+forgotten-password flows respond. The positional argument takes more than one address, which is how
+you run a whole normalisation set in a single pass.
```bash
-pipx install holehe
+pipx install holehe # or: pip3 install holehe
+
holehe target@example.com
+
+# several spellings at once -- the positional argument is variadic
+holehe d.whitfield@example.com dwhitfield@example.com dwhitfield+shop@example.com
+
+# only the hits, which is the only readable form on a 120-site run
holehe target@example.com --only-used
+
+# skip the password-recovery methods: fewer sites checked, much less chance of mail
+holehe target@example.com -NP
+
+# keep the run: writes holehe_<timestamp>_<email>_results.csv in the working directory
+holehe target@example.com -C
+
+# slow or rate-limiting sites, and output that survives a pipe
+holehe target@example.com -T 20 --no-color --no-clear
+```
+
+The output is one line per site, prefixed `[+]` used, `[-]` not used and `[x]` rate-limited. That
+is the whole vocabulary — holehe's own legend prints those three and nothing else, so a module that
+threw an exception is folded into one of them rather than flagged, and you cannot tell a broken
+check from a clean negative by reading the terminal. Where a site leaks more than existence,
+holehe appends it to the same line — a partially
+masked recovery email, a masked phone number, sometimes a full name or an account creation date.
+Those appended fragments are the most valuable thing the tool produces and the easiest to lose in
+the scroll, which is the argument for `-C` and a look at the CSV rather than the terminal.
+
+How it misleads you: upstream states the tool does not alert the target, and for most of the 120
+sites that holds, but the underlying mechanism is the forgotten-password function — treat
+"no notification" as a claim about the common case and not a guarantee, and reach for `-NP` when a
+single unexpected reset email would end the engagement. Rate limiting is per egress IP and
+accumulates across the run, so the sites checked last are the most likely to report `[x]`; a
+second run from a different address will produce a different-looking result set from the same
+truth. Never fold `[x]` into `[-]` when you write up.
+
+### Zehef
+
+A second opinion with a different module mix, useful because its overlap with holehe is partial.
+It runs paste-site and breach-index checks alongside account checks, and generates candidate
+address combinations from a name.
+
+```bash
+git clone https://github.com/N0rz3/Zehef.git
+cd ./Zehef
+pip3 install -r requirements.txt
+
+python3 zehef.py target@example.com
+python3 zehef.py -h
```
-**This one is not passive.** Some providers email the address to say a reset was attempted. Know
-that before you run it on a live target.
+The combination generator is the part to be careful with. It emits plausible address spellings for
+a person, which is a legitimate way to build a candidate list for the oracles above — and nothing
+it emits is a confirmed address. Keep generated strings in a separate column from observed ones; a
+generated address that later produces a `[+]` somewhere is a finding, and a generated address in a
+report without that step is a fabrication.
+
+### Syntax, DNS and SMTP
+
+Before any tool, ask whether the domain can receive mail at all. This is two DNS queries, it is
+invisible to everybody, and it regularly ends the investigation early.
+
+```bash
+# the mail exchangers for the domain
+dig +noall +answer MX example.com
+# example.com. 126 IN MX 0 .
+
+# SPF, DMARC and the rest of the policy record set
+dig +short TXT example.com
+# "v=spf1 -all"
+dig +short TXT _dmarc.example.com
+```
-### Verification and syntax
+A single MX of `0 .` is the null MX of RFC 7505: the domain is declaring that it accepts no mail
+at all, so no mailbox on it exists, nothing was ever delivered to it and no account can have been
+confirmed through it. Paired with `v=spf1 -all` — the domain sends no mail either — it means the
+address in front of you is decorative. The opposite finding is just as useful: MX records pointing
+at `*.l.google.com` or `*.outlook.com` tell you a custom-domain address is really a Google or
+Microsoft account, which puts GHunt and the Microsoft modules back on the table for an address that
+does not look like a free-mail one.
-Use a validator to confirm deliverability and to spot disposable domains before you invest effort.
-`metadata2go`, Email Checker and similar services cover this. An MX lookup on the domain is the
-quick manual version:
+Confirming a specific mailbox, rather than the domain, means an SMTP conversation, and that is an
+active step with a log entry at the other end.
```bash
-dig +short MX example.com
+# RCPT probe: open the transaction, offer the recipient, quit before DATA.
+# Nothing is sent. The server's reply to RCPT TO is the entire result.
+swaks --to d.whitfield@example.com --from '<>' --quit-after RCPT --timeout 10
+
+# client-side commands only, which is the readable form for a transcript
+swaks --to d.whitfield@example.com --from '<>' --quit-after RCPT \
+ --no-info-hints --hide-receive --hide-informational
+
+# name the server explicitly when you want to probe one specific MX
+swaks --to d.whitfield@example.com --from '<>' --quit-after RCPT --server mx1.example.net
```
+The trap is the first line of output. Swaks resolves the recipient domain's MX itself using Perl's
+`Net::DNS`, and when that module is missing it says so and falls back to localhost:
+
+```text
+*** MX Routing not available: requires Net::DNS. Using localhost as mail server
+=== Trying localhost:25...
+*** Error connecting to localhost:25:
+*** Connection refused
+```
+
+That is not a result about the target. Install `Net::DNS` or pass `--server` from the `dig` output
+above, and read the banner before you read the reply. Beyond that: a catch-all domain accepts every
+recipient, so a `250` there means nothing; greylisting and tarpitting produce temporary rejections
+that look like refusals; and large providers long ago stopped answering this question honestly.
+Enumerating mailboxes through an authentication or delivery endpoint is also the step most likely
+to be regulated where you are standing, so settle that before you type it, not after. DNS, WHOIS
+and certificate-transparency work on the domain itself belongs on
+[Websites & Infrastructure](/sheets/osint/websites-and-infrastructure).
+
## Phone
+### phonenumbers (libphonenumber)
+
+The Python port of Google's libphonenumber, and the correct first step for every number you touch.
+It is offline, instant, free, and it answers the structural questions before you spend money or
+attention on anything else.
+
+```bash
+pip install phonenumbers
+```
+
+```python
+import phonenumbers
+from phonenumbers import carrier, geocoder, timezone, PhoneNumberFormat, PhoneNumberType
+
+# parse() needs a region hint for anything not in +E.164 form, and raises
+# NumberParseException rather than guessing:
+# phonenumbers.parse("020 7946 0123", None)
+# -> NumberParseException: (0) Missing or invalid default region.
+n = phonenumbers.parse("+447911123456", None)
+
+print(phonenumbers.format_number(n, PhoneNumberFormat.E164)) # +447911123456
+print(phonenumbers.format_number(n, PhoneNumberFormat.INTERNATIONAL)) # +44 7911 123456
+print(phonenumbers.format_number(n, PhoneNumberFormat.NATIONAL)) # 07911 123456
+
+print(phonenumbers.is_possible_number(n)) # True
+print(phonenumbers.is_valid_number(n)) # True
+print(phonenumbers.number_type(n) == PhoneNumberType.MOBILE) # True
+
+print(carrier.name_for_number(n, "en")) # JT
+print(geocoder.description_for_number(n, "en")) # Guernsey
+print(timezone.time_zones_for_number(n))
+# ('Europe/Guernsey', 'Europe/Isle_of_Man', 'Europe/Jersey', 'Europe/London')
+
+print(phonenumbers.region_code_for_number(n)) # GG
+print(phonenumbers.is_valid_number_for_region(n, "GB")) # False
+print(phonenumbers.is_mobile_number_portable_region("GB")) # True
+
+# why a number failed, rather than just that it did
+z = phonenumbers.parse("+44770090012345", None)
+print(phonenumbers.is_possible_number_with_reason(z)) # 3 (TOO_LONG)
+
+# a known-good number for a region, to sanity-check your own parsing
+print(phonenumbers.format_number(phonenumbers.example_number("GB"),
+ PhoneNumberFormat.E164)) # +441212345678
+```
+
+That output contains the lesson. A number that looks like an ordinary UK mobile resolves to region
+`GG`, carrier `JT` and four candidate timezones, because `+44` is shared between the United
+Kingdom, Guernsey, Jersey and the Isle of Man. If you had written "UK subscriber" from the `+44`
+you would have the wrong jurisdiction, the wrong regulator and the wrong carrier to serve anything
+on. `number_type` returns `PhoneNumberType.UNKNOWN`, which is `99` and not `0` — `0` is
+`FIXED_LINE`, so a truthiness check on the result is silently wrong.
+
+How it misleads you: the carrier field is the *original* carrier. Upstream is explicit that for
+regions supporting portability "we return the original carrier for the supplied number", so in any
+country with mobile number portability — which `is_mobile_number_portable_region` will tell you
+about — the carrier name is a historical fact about a block, not a current fact about a subscriber.
+An empty string means the mapping has no data, not that the number has no carrier. The geocoder
+returns a numbering-plan area, which for a fixed line is roughly where the line is and for a mobile
+is roughly where the block was issued; neither is a residence. And none of it speaks to whether
+the number is in service.
+
### PhoneInfoga
-Format normalisation, carrier and line-type identification, and footprint searching for a number.
+Format normalisation, carrier and line-type identification, and footprint searching, wrapped in a
+CLI and a local REST API. It is a Go binary: the documented installs are the release archive,
+Homebrew and Docker, and there is no pip package to reach for.
```bash
-pipx install phoneinfoga
+brew install phoneinfoga
+# or: bash <( curl -sSL https://raw.githubusercontent.com/sundowndev/phoneinfoga/master/support/scripts/install )
+# sudo install ./phoneinfoga /usr/local/bin/phoneinfoga
+# or: docker pull sundowndev/phoneinfoga:latest
+
+phoneinfoga version
phoneinfoga scan -n "+14155550123"
-phoneinfoga serve # local web UI
+phoneinfoga scan -n "+1 (555) 444-1212" # ( ) - + are escaped for you
+
+# the local web UI and REST API on port 5000, or a port you choose
+phoneinfoga serve
+phoneinfoga serve -p 8080
+phoneinfoga serve --no-client # REST API only, no web client
+
+# a compiled custom scanner
+phoneinfoga scan -n "+14155550123" --plugin ./custom_scanner.so
+
+# docker equivalents
+docker run --rm -it sundowndev/phoneinfoga version
+docker run --rm -it -p 5000:5000 sundowndev/phoneinfoga serve
+```
+
+Five scanners ship with it — Local, Numverify, Googlesearch, Googlecse and OVH — and they are
+configured by environment variable, not by flag:
+
+```bash
+export NUMVERIFY_API_KEY=... # Numverify scanner
+export GOOGLE_API_KEY=... # Googlecse scanner
+export GOOGLECSE_CX=... # Googlecse search engine ID
+export GOOGLECSE_MAX_RESULTS=50 # optional, default 10, maximum 100
+```
+
+Local and Googlesearch and OVH need no configuration, which means an unconfigured install silently
+runs a subset: a scan with no `NUMVERIFY_API_KEY` set is not a scan that found no carrier. Read the
+country code as mandatory — the docs say it plainly — because a national-format number scanned
+without one produces confident output about the wrong country. The Googlesearch scanner's value is
+a list of dork URLs you still have to open and judge yourself; treat it as a worklist, not a
+finding. And the local scanner is numbering-plan metadata, so it inherits every portability caveat
+from the section above.
+
+### Phunter
+
+A number-oriented sweep with a reverse-lookup module and an Amazon account check, which is a
+different and narrower bet than PhoneInfoga's.
+
+```bash
+git clone https://github.com/N0rz3/Phunter.git
+cd ./Phunter
+pip3 install -r requirements.txt
+
+python3 phunter.py -t "+33644637111" # operator, location, line type, reputation
+python3 phunter.py -f numbers.txt # a file of numbers
+python3 phunter.py -a "+33644637111" # is the number attached to an Amazon account
+python3 phunter.py -p "+33644637111" # reverse lookup: who the number belongs to
+python3 phunter.py -p "+33644637111" -o owner.txt
+python3 phunter.py -v # version and service status
```
+The `-v` check is worth running first on a tool of this kind: these modules break when a target
+site changes its flow, and a module that is quietly broken returns the same empty result as a
+module that found nothing. `-a` and `-p` are the two that matter and they are the two that reach
+out — the Amazon check exercises an account flow, and the reverse lookup queries third-party
+directories whose coverage collapses outside a handful of countries. A "reputation" or spam flag
+here describes crowd-sourced complaints about the number, which is evidence about the number's
+behaviour and not about its owner.
+
+### ignorant
+
+The phone-side counterpart to holehe, by the same author and with the same flag vocabulary. It
+checks whether a number is registered on a short list of sites — the README names Snapchat and
+Instagram — and the positional arguments are the part people get wrong.
+
+```bash
+pip3 install ignorant
+
+# country code WITHOUT the plus, then the national number WITHOUT the trunk zero
+ignorant 33 644637111 # = +33 6 44 63 71 11
+ignorant 44 7911123456 # = +44 7911 123456, not "07911 123456"
+
+ignorant 33 644637111 --only-used
+ignorant 33 644637111 -T 20 --no-color --no-clear
+```
+
+Same prefixes as holehe: `[+]` registered, `[-]` not registered, `[x]` rate-limited. Same
+discipline applies — it works by poking registration flows, so it is active, and a short result
+set from a long run usually means you were throttled rather than that the number is unused. Because
+the site list is small, a clean sweep of negatives is close to worthless on its own; its use is
+confirmation when you already have a candidate and want a second thread tying it to an account.
+
+### Telegram Phone Number Checker
+
+Bellingcat's tool for the single most productive messenger check, because Telegram will confirm
+registration and often hand over a username, a display name and a user ID.
+
+```bash
+pip install telegram-phone-number-checker
+# with proxy support: pip install telegram-phone-number-checker[proxy]
+
+# credentials come from https://my.telegram.org/ and live in a .env file,
+# or on the command line
+telegram-phone-number-checker --phone-numbers "+447911123456,+33644637111"
+
+telegram-phone-number-checker --phone-numbers "+447911123456" \
+ --download-profile-photos --output whitfield.json
+
+telegram-phone-number-checker --usernames "someuser,otheruser"
+
+telegram-phone-number-checker --phone-numbers "+447911123456" \
+ --api-id 123456 --api-hash 0123456789abcdef0123456789abcdef \
+ --api-phone-number "+447700900123"
+
+telegram-phone-number-checker --phone-numbers "+447911123456" --proxy socks5://127.0.0.1:9050
+```
+
+The user ID in the output is the durable key — usernames and display names change weekly, the ID
+does not — and the profile photo is a direct pivot to
+[reverse image search](/sheets/osint/reverse-image-search) and to
+[Social Media Platforms](/sheets/osint/social-media-platforms) for the rest of the account.
+
+How it misleads you: the lookup runs as your own Telegram account, and upstream advises plainly
+"not to use your personal account for automations as telegram may block it". A result of
+"no username detected" is a privacy setting, not an absence of an account, and an error line means
+the account exists but has restricted phone-number discovery — which is itself a finding about how
+careful the owner is. Default `--output` is `results.json` in the working directory, so name it per
+target or you will overwrite the last run.
+
### Caller-ID and directory databases
Crowd-sourced caller-ID apps hold name data for numbers that appear nowhere else, because their
@@ -115,24 +622,166 @@ users uploaded their own address books.
| [NigeriaPhonebook](https://nigeriaphonebook.com/) | Nigerian number-to-name lookups. | free |
Searching a number in these apps can be visible to other users of the same app, and uploading a
-contact list to get access hands over everyone in your phone. Use a clean device.
+contact list to get access hands over everyone in your phone. Use a clean device. Read a hit as
+"at least one stranger saved this number under this label at some point", which is a real signal
+and is not identification — the same number commonly returns a personal name, a trade name and an
+insult, all from different users, all years apart.
### Messaging-app checks
Many messengers confirm whether a number is registered, and often expose a profile photo and
status. Adding the number as a contact may make you visible to them — check the platform's
-behaviour before you do it on a target who might notice.
+behaviour before you do it on a target who might notice. Do these from a dedicated device and
+number: the contact list on the account you use for checks is itself a disclosure, and some
+platforms surface "contacts who joined" to the people you looked up.
+
+## Identifier tools at a glance
+
+| Tool | Identifier | What it does | Touches the target |
+| --- | --- | --- | --- |
+| [Epieos](https://epieos.com/) | email, phone | One server-side query across Google, Microsoft, Gravatar and consumer services. | no |
+| [GHunt](https://github.com/mxrch/GHunt) | email, Gaia ID | Google account depth: internal ID, photo, Maps, calendar, Drive. | no, but runs as your Google session |
+| Gravatar hash probe | email | Deterministic SHA-256 lookup for an avatar and a linked profile. | no |
+| `dig` MX/TXT | email domain | Whether the domain can receive mail at all, and who runs it. | no |
+| [holehe](https://github.com/megadose/holehe) | email | Registration and reset-flow checks across 120+ sites. | yes — may generate mail |
+| [Zehef](https://github.com/N0rz3/Zehef) | email | Paste and breach checks, account checks, address generation. | partly |
+| [swaks](https://www.jetmore.org/john/code/swaks/) | email | SMTP RCPT probe against the real mail exchanger. | yes — logged at the server |
+| [phonenumbers](https://github.com/daviddrysdale/python-phonenumbers) | phone | Offline parse, validity, type, carrier block, timezone. | no |
+| [PhoneInfoga](https://github.com/sundowndev/phoneinfoga) | phone | Normalisation plus scanner modules and a local REST API. | no for Local, yes for footprint scanners |
+| [Phunter](https://github.com/N0rz3/Phunter) | phone | Line metadata, reverse lookup, Amazon account check. | yes for `-a` and `-p` |
+| [ignorant](https://github.com/megadose/ignorant) | phone | Registration checks on a small set of sites. | yes |
+| [Telegram Phone Number Checker](https://github.com/bellingcat/telegram-phone-number-checker) | phone | Telegram registration, username, display name, user ID. | yes — runs as your account |
## Pitfalls
-- **Reset-flow probing is active.** holehe and similar tools can generate mail to the target.
- Passive-only work means Epieos, GHunt and string searching, nothing that touches a login flow.
+- **"Valid" does not mean "assigned".** `+1 415 555 0123` validates cleanly and is reserved for
+ fiction. Validity is a statement about the numbering plan, nothing more; service status costs
+ money or a ring.
+- **Possible is weaker than valid, and both are weaker than real.** A number can be possible,
+ invalid and still in a dataset as a deliberate placeholder.
+- **`+44` is not the United Kingdom.** It is shared with Guernsey, Jersey and the Isle of Man, and
+ `region_code_for_number` will tell you which — after which your jurisdiction, regulator and
+ carrier are all different.
+- **Carrier data is pre-portability.** In a portable region the lookup returns the carrier the
+ block was issued to, which may be two providers out of date. An empty carrier field means no
+ data, not no carrier.
+- **Reset-flow probing is active.** holehe, ignorant, Phunter's `-a` and any SMTP probe can
+ generate mail or a log entry. Passive-only work means Epieos, GHunt, DNS, hash probes and string
+ searching, nothing that touches a login flow.
+- **`[x]` is not `[-]`.** Rate-limited is not absent. Rate limits are per egress IP and get worse
+ through a run, so the sites checked last are the least trustworthy in the output, and a second
+ run from elsewhere will disagree with the first.
+- **Run a control address.** Enumeration-hardened sites return the same friendly page for every
+ input. Without a known-nonexistent control in the same run, you cannot tell that apart from a
+ negative.
+- **Null MX is a finding, not an error.** An MX of `0 .` means the domain accepts no mail, so the
+ address never worked. Catch-all domains are the mirror image: they accept everything, so
+ "the address exists" is meaningless on them.
+- **One mailbox, many keys.** Dots and `+tags` are the same Gmail inbox and completely different
+ hashes, dataset rows and tool inputs. Normalise before you count, and probe every spelling
+ before you record an absence.
+- **A masked tail is a 1-in-100 coincidence.** Two digits of a recovery number matching your
+ candidate is weak evidence. Two independent masks from different services multiplying out to
+ 1 in 100,000 is strong. Say which one you have.
+- **A recovery phone is not the subject's phone.** It is frequently a partner's, a parent's, a
+ former employer's desk line or a long-dead PAYG handset kept for exactly this purpose.
- **Caller-ID names are user-submitted.** "John Smith Plumber" may be what one stranger saved the
- number as years ago.
-- **Recycled numbers.** Carriers reissue numbers after a few months of disuse, so an old
- registration may belong to someone unrelated.
-- **Catch-all domains** accept every address, so "the address exists" is meaningless on them.
-- **Aggregated identifiers age badly.** An address that reached someone in 2019 may be dead now.
+ number as years ago, and the lookup itself may be visible to that app's other users.
+- **Recycled numbers and recycled mailboxes.** Carriers reissue numbers after a few months of
+ disuse and some providers release abandoned usernames, so an old registration may belong to
+ someone unrelated — and account-existence output will not distinguish them.
+- **Aggregated identifiers age badly.** An address that reached someone in 2019 may be dead now;
+ date every lookup and quote the date beside the claim.
+- **Enumeration may be illegal where you are.** Probing authentication endpoints for account
+ existence is regulated in some jurisdictions regardless of intent. Settle that before the first
+ command, not in the write-up.
+
+## Worked example
+
+One datum: a tip-off naming **Dana Whitfield**, with one email address and one phone number and no
+other detail — `d.whitfield@example.com` and `+44 7700 900123`. The task is to decide whether this
+record is worth an investigation before anyone spends a day on it.
+
+1. **Parse the number offline, first.** No network, no cost, no trace.
+
+```text
+input : +447700900123
+E164 : +447700900123
+INTERNATIONAL : +44 7700 900123
+NATIONAL : 07700 900123
+country_code : 44 national_number: 7700900123
+is_possible : True
+is_valid : False
+number_type : 99 (PhoneNumberType.UNKNOWN; MOBILE would be 1)
+carrier : ''
+geocoder : ''
+timezones : ('Etc/Unknown',)
+region : None
+```
+
+2. **Read the disagreement.** Possible but not valid, with no region, no carrier and
+ `Etc/Unknown` for a timezone. Twelve digits in the right shape for `+44`, matching no block the
+ numbering plan defines. That pattern is a reserved or unallocated range, not a typo.
+3. **Name the range.** The national number is `7700900123`, which sits inside `07700 900000` to
+ `07700 900999` — Ofcom's drama block, 1,000 numbers that Ofcom states "will not be allocated to
+ communications providers in the foreseeable future". The number is fictional by construction.
+4. **Do not take a tool's word for the conclusion.** Running a registration oracle against it
+ returns negatives, and those negatives are worthless: a rate limit, a privacy setting and an
+ unassigned number all print the same thing. The load-bearing evidence is the published range,
+ which anyone can check in a browser, not a `[-]` in a terminal.
+5. **Switch to the address and ask the cheap question.** Two DNS queries, invisible to everyone:
+
+```text
+$ dig +noall +answer MX example.com
+example.com. 126 IN MX 0 .
+
+$ dig +short TXT example.com | grep spf
+"v=spf1 -all"
+```
+
+The `grep` is not decoration. A bare `TXT` query on a live domain also returns whatever
+domain-verification tokens are parked there that week, and those rotate, so filtering is what makes
+the transcript something a reader can reproduce rather than something they have to trust. The MX
+TTL will differ from the one above for the same reason.
+
+6. **Read it.** A single MX of `0 .` is RFC 7505's null MX — the domain publishes that it accepts
+ no mail — and `v=spf1 -all` says it sends none either. No mailbox on that domain has ever
+ received anything, so no service can have confirmed an account through it.
+7. **Confirm with the deterministic probe, on every spelling.** SHA-256 of the trimmed, lower-cased
+ address, against the avatar endpoint with `d=404`:
+
+```text
+d.whitfield@example.com fc9518a0411a19b071630d4b153e4774e0697053e4015e5eb9e1ab5c934021a3 HTTP 404
+dwhitfield@example.com 6a2ae976758ac8100467b0bea363301fcb6fe756c8262b95336be1d03dcbd919 HTTP 404
+dwhitfield+shop@example.com bbb5413ad03bc64773f79b7804cd4868f34e7fd7a5fa2e1808617ed45289369f HTTP 404
+```
+
+8. **Total cost.** Four lookups, no packet sent to any mail server, nothing delivered to anybody,
+ about ninety seconds. Both identifiers in the tip-off come from reserved or
+ non-deliverable space, and they come from two different reserved spaces, which is the signature
+ of a synthetic record rather than a transcription error.
+9. **What the same pipeline looks like on a live record,** for contrast: `+44 7911 123456` parses
+ as valid, `PhoneNumberType.MOBILE`, carrier `JT`, geocoder `Guernsey`, region `GG`. That is the
+ point at which the oracles are worth running — and the point at which you must notice that your
+ `+44` number is not a GB number at all.
+
+What you can assert: that the phone number lies inside a block Ofcom publishes as reserved for
+drama and unallocated, and that the email domain publishes a null MX and an SPF record refusing all
+sending. Both claims are independently checkable in under a minute by anyone reading the report,
+and together they say the record is synthetic. The names, dates and timestamps of the two DNS
+queries go in the notes, because DNS is mutable and a null MX today is not a null MX last year.
+
+What would falsify it: Ofcom reallocating the drama range, which it says it will not do but which
+is a policy rather than a law of nature; a null MX that was added after the record was created,
+which would mean the address had worked earlier and is the single most likely way this conclusion
+is wrong; or the tip-off being a mangled transcription of a real number one digit away from this
+one. The weakest measurement is the Gravatar probe — a 404 speaks only for the exact strings
+hashed, and `d.whitfield` and `dwhitfield` produce entirely unrelated digests, so three 404s cover
+three spellings and not a mailbox. Quote it as "no avatar for these three spellings", never as
+"no account". And none of this says anything at all about whether a person called Dana Whitfield
+exists; it says the two identifiers attached to the name do not. Converting a name into a person is
+[people search](/sheets/osint/people-search), and converting a handle into one is
+[Usernames & Accounts](/sheets/osint/usernames-and-accounts).
## Broader catalogues
diff --git a/src/content/sheets/osint/image-video-forensics.md b/src/content/sheets/osint/image-video-forensics.md
@@ -438,12 +438,17 @@ the same shopfront in a news photo dated fourteen months earlier, and
[geolocation](/sheets/osint/geolocation) puts the corner two streets from the one named. Shadow
direction in the frame is consistent with late afternoon, not morning.
-What you can defensibly say: the file was produced by ffmpeg, contains at least three edits, is at
+What you can assert: the file was produced by ffmpeg, contains at least three edits, is at
least fourteen months old in part, and shows a different corner than claimed — each tied to a named
command, its output, and the hash `9f2c...e41b`. What you cannot say from the file alone is who
assembled it or why. That needs the second independent source: the earlier news photo, which is
what actually carries the date.
+What would falsify it: the earlier image proving to be a later backdated copy, the scene-change
+threshold firing on flashes rather than edits, or the analysed hash not matching the preserved
+file. The fourteen-month bound rests on the independently dated earlier image; the ffmpeg encoder
+and frame-rate evidence establish re-encoding, not the age of the depicted footage.
+
## Broader catalogues
- [Image and Video Analysis OSINT](https://tools.osintnewsletter.com/tool-categories/image-and-video-analysis-osint)
diff --git a/src/content/sheets/osint/maps-and-satellite-imagery.md b/src/content/sheets/osint/maps-and-satellite-imagery.md
@@ -4,9 +4,9 @@ description: "Choose the right imagery source for the question, find historical
category: osint
subcategory: "Geospatial"
tags: [osint, satellite, maps, gis, street-view]
-tools: [qgis, gdal, overpass, copernicus-data-space, google-earth-pro, mapillary, earthexplorer]
+tools: [qgis, gdal, ogr2ogr, overpass, copernicus-data-space, google-earth-pro, google-earth-engine, mapillary, earthexplorer, landsatlook, stac-client, nasa-firms, nasa-gibs]
difficulty: intermediate
-updated: 2026-09-28
+updated: 2026-10-04
references:
- name: "Bellingcat's Online Investigation Toolkit"
url: "https://bellingcat.gitbook.io/toolkit"
@@ -20,6 +20,36 @@ references:
license: none
relation: derived
note: "Second tool catalogue, cross-checked against the above."
+ - name: "GDAL programs documentation"
+ url: "https://gdal.org/en/stable/programs/"
+ author: "OSGeo"
+ relation: link-only
+ note: "Upstream reference for every gdal* and ogr* flag quoted on this page."
+ - name: "Overpass QL reference"
+ url: "https://wiki.openstreetmap.org/wiki/Overpass_API/Overpass_QL"
+ author: "OpenStreetMap contributors"
+ relation: link-only
+ note: "Upstream reference for the Overpass settings, filters and out modes used here."
+ - name: "Copernicus Data Space OData API"
+ url: "https://documentation.dataspace.copernicus.eu/APIs/OData.html"
+ author: "Copernicus Data Space Ecosystem"
+ relation: link-only
+ note: "Upstream reference for the $filter, $orderby and attribute query syntax."
+ - name: "Mapillary API documentation"
+ url: "https://www.mapillary.com/developer/api-documentation"
+ author: "Mapillary"
+ relation: link-only
+ note: "Upstream reference for the Graph API endpoints, parameters and image fields."
+ - name: "NASA GIBS API documentation"
+ url: "https://nasa-gibs.github.io/gibs-api-docs/"
+ author: "NASA EOSDIS"
+ relation: link-only
+ note: "Upstream reference for the WMTS and WMS endpoint templates."
+ - name: "NASA FIRMS Area API"
+ url: "https://firms.modaps.eosdis.nasa.gov/api/area/"
+ author: "NASA FIRMS"
+ relation: link-only
+ note: "Upstream reference for the fire-detection API path parameters."
---
## What this covers
@@ -52,7 +82,7 @@ depth trade off against each other and no single provider wins on all three.
| What does this place look like in detail? | Google Earth Pro, Esri World Imagery — sub-metre, but infrequent |
| What changed between two dates? | Sentinel-2 (5-day revisit, 10m), Landsat (16-day, 30m, back to 1972) |
| What did it look like years ago? | Google Earth Pro's historical slider, Landsat archive |
-| Was there a fire / flood / new construction? | Sentinel-2 false-colour composites |
+| Was there a fire / flood / new construction? | Sentinel-2 false-colour composites, NASA FIRMS |
| What is at street level? | Google Street View, Mapillary, KartaView, Yandex Panoramas |
| What features exist here, as data? | OpenStreetMap via Overpass |
@@ -60,6 +90,29 @@ Free high-cadence optical imagery bottoms out around 10m per pixel. Anything fin
and usually costs real money, so plan around Sentinel for change detection and reserve high-res for
confirming a specific thing.
+## Imagery sources at a glance
+
+Resolution is the number people quote and the least useful of the four. Revisit and archive depth
+decide whether a question is answerable at all, and scriptability decides whether you can answer it
+for fifty coordinates instead of one.
+
+| Source | Resolution | Revisit | Archive back to | Scriptable access |
+| --- | --- | --- | --- | --- |
+| Sentinel-2 (Copernicus) | 10m visible | ~5 days | 2015 | OData and STAC, free, token for download |
+| Sentinel-1 (radar) | 5x20m | ~6 days | 2014 | same OData catalogue |
+| Landsat 8/9 | 30m, 15m pan | 16 days each | 1972 across the series | LandsatLook STAC, no key for search |
+| MODIS / VIIRS | 250m–1km | sub-daily | 2000 / 2012 | GIBS WMTS and WMS, no key |
+| Google Earth Pro | sub-metre | irregular, years apart | varies by place, often 1985 | none; desktop export only |
+| Esri World Imagery | sub-metre | irregular | current only, no slider | XYZ tile URL |
+| Mapillary | street level | contributor-driven | 2014 | Graph API, free token |
+| OpenStreetMap | vector, not imagery | continuous | full edit history | Overpass API, no key |
+
+Two consequences worth internalising. Nothing free gives you both sub-metre resolution and a dated
+archive, which is why high-resolution work means Google Earth Pro screenshots with the status-bar
+date, and change detection means Sentinel. And the published revisit figures are orbital, not
+usable: see the cloud arithmetic in the worked example below, where a nominal five-day revisit
+yielded eighteen scenes in a month and two worth opening.
+
## Historical imagery
**Google Earth Pro** is free desktop software and its historical imagery slider is the most
@@ -67,7 +120,8 @@ accessible archive of high-resolution coverage. Note the imagery date shown at t
the single most important piece of context and the most commonly ignored.
**Landsat** goes back to 1972 and is the only free option for multi-decade change. Browse it via
-[EarthExplorer](https://earthexplorer.usgs.gov/), covered below.
+[EarthExplorer](https://earthexplorer.usgs.gov/), covered below, or query the same archive through
+the LandsatLook STAC API, which needs no account.
## Street-level beyond Google
@@ -124,6 +178,65 @@ curl -s -A "$UA" "$OVERPASS" --data-urlencode \
-o aerodromes.json
```
+Four forms of the query that earn their keep once the search is more than one box. `nwr` matches
+nodes, ways and relations in one pass, which matters because a feature mapped as a node in one
+country is a way in the next; a `node[...]` query silently misses half of them. The `[bbox:...]`
+setting applies to every statement that carries no explicit box, so an exploratory query stops
+being a wall of repeated coordinates. `out:csv` with named fields lands in `awk` or a spreadsheet
+without a `jq` filter in between. And `(newer:...)` asks OSM's own edit history what changed,
+which is a change-detection signal that costs nothing and arrives before any satellite pass.
+
+```bash
+# CSV with chosen fields: header line on, comma separator. Note the renaming --
+# ::id comes back as the column @id, which is what you grep for afterwards
+curl -s -A "$UA" "$OVERPASS" --data-urlencode \
+ 'data=[out:csv(::id,::type,::lat,::lon,"name";true;",")][timeout:30];
+ node["amenity"="pharmacy"](52.37,4.88,52.38,4.90);out center;'
+# @id,@type,@lat,@lon,name
+# 1819064252,node,52.3723707,4.8940844,Dam Apotheek
+# 2720875314,node,52.3783767,4.8823525,Medicijnman Apotheek Jordaan
+
+# nwr, so a feature mapped as a way somewhere is not missed. Counting first:
+curl -s -A "$UA" "$OVERPASS" --data-urlencode \
+ 'data=[out:csv(::count)][timeout:60];area["ISO3166-1"="NL"]->.a;
+ nwr["amenity"="fuel"](area.a);out count;'
+# @count
+# 4117
+
+# a global bbox in the settings, so each statement inherits it
+curl -s -A "$UA" "$OVERPASS" --data-urlencode \
+ 'data=[bbox:52.3730,4.8900,52.3740,4.8915][out:json][timeout:25];nwr["historic"];out center;'
+
+# what OSM itself says changed: buildings touched since a date, with edit metadata
+curl -s -A "$UA" "$OVERPASS" --data-urlencode \
+ 'data=[out:json][timeout:60];
+ way["building"](52.3730,4.8900,52.3740,4.8915)(newer:"2020-01-01T00:00:00Z");out meta center;' \
+ | jq -r '.elements[] | "\(.timestamp) v\(.version) \(.id)"'
+```
+
+Getting the result into a GIS is the step people do by hand and should not. Overpass will emit OSM
+XML, which GDAL's OSM driver reads directly, so one `curl` and one `ogr2ogr` produce a GeoPackage.
+The recurse-down operator `>` is what makes the XML usable: without `(._;>;)` you get ways with no
+node coordinates and every geometry comes out empty.
+
+```bash
+# ways plus the nodes that define them, as OSM XML
+curl -s -A "$UA" "$OVERPASS" --data-urlencode \
+ 'data=[out:xml][timeout:60];way["building"](52.3730,4.8900,52.3740,4.8915);(._;>;);out meta;' \
+ -o area.osm
+
+# the driver exposes five fixed layers: points, lines, multilinestrings,
+# multipolygons, other_relations. Closed building ways land in multipolygons.
+ogr2ogr -f GPKG area.gpkg area.osm multipolygons -nln buildings
+
+# tag filtering needs the SQLite dialect, because tags arrive as one other_tags blob
+ogr2ogr -f GeoJSON masts.geojson area.osm -dialect sqlite \
+ -sql "SELECT * FROM points WHERE other_tags LIKE '%\"man_made\"=>\"mast\"%'"
+
+# reproject on the way out, so the areas and distances you measure are in metres
+ogr2ogr -f GPKG -t_srs EPSG:32631 buildings_utm.gpkg area.gpkg buildings
+```
+
Queries are metered by server CPU time, not request count: a careless `[timeout:900]` over a whole
country will get you a 429 and then a temporary ban from the main instance. Run the count form
first. Because OSM is crowd-sourced, a feature's absence means nobody mapped it — never that it
@@ -142,12 +255,21 @@ photograph or scanned map, and build band maths like NDVI without opening a GUI.
# what is in this file: CRS, extent, pixel size, bands, and the capture metadata
gdalinfo -stats scene.tif | head -40
+# the same, machine-readable, when you are checking fifty files rather than reading one
+gdalinfo -json scene.tif | jq '{crs: .coordinateSystem.wkt[0:60], size, bands: [.bands[].band]}'
+
# crop to a bounding box in the file's own CRS, so you stop moving 2GB around
gdal_translate -projwin 4.88 52.38 4.90 52.37 -projwin_srs EPSG:4326 scene.tif crop.tif
+# the same crop by pixel window, when you are working off a screenshot's coordinates
+gdal_translate -srcwin 2400 1850 512 512 scene.tif tile.tif
+
# reproject to Web Mercator for overlay on a slippy basemap
gdalwarp -t_srs EPSG:3857 -r cubic crop.tif crop_3857.tif
+# clip to an area of interest polygon rather than a rectangle, and trim the canvas to it
+gdalwarp -cutline aoi.gpkg -crop_to_cutline -dstalpha -overwrite scene.tif aoi_only.tif
+
# georeference a scanned map: four ground control points (pixel x, pixel y, lon, lat), then warp
gdal_translate -of GTiff -a_srs EPSG:4326 \
-gcp 120 95 4.8855 52.3805 -gcp 1890 110 4.9015 52.3799 \
@@ -162,18 +284,44 @@ gdalbuildvrt mosaic.vrt tiles/*.tif
gdal_calc.py -A B08.jp2 -B B04.jp2 --outfile=ndvi.tif \
--calc="(A.astype(float)-B)/(A.astype(float)+B+0.0001)"
+# differencing two dates: --extent=intersect is what stops a silent misalignment
+gdal_calc.py -A ndvi_may.tif -B ndvi_sep.tif --outfile=ndvi_delta.tif \
+ --calc="B-A" --type=Float32 --extent=intersect --projectionCheck --overwrite
+
# hillshade from a DEM, for reading terrain in a photograph's background
gdaldem hillshade -z 2 dem.tif hillshade.tif
+# multidirectional hillshade keeps slopes facing away from the light readable
+gdaldem hillshade -multidirectional -compute_edges dem.tif hillshade_multi.tif
+
+# slope in degrees, for arguing about whether a vehicle track is plausible
+gdaldem slope -compute_edges dem.tif slope.tif
+
# a shareable PNG at a sane size, with the world file so it stays georeferenced
gdal_translate -of PNG -outsize 25% 25% -co WORLDFILE=YES crop.tif preview.png
```
+The one command that settles arguments is `gdallocationinfo`: it reads the pixel value at a
+coordinate, which turns "that looks darker" into a number you can put in a report.
+
+```bash
+# the band values under one WGS84 coordinate, values only, coordinate echoed back
+echo "4.8909 52.3738" | gdallocationinfo -wgs84 -valonly -E -field_sep , scene.tif
+
+# a whole candidate list in one pass: stdin is read line by line
+gdallocationinfo -wgs84 -valonly -E -field_sep , ndvi_delta.tif < candidates.txt
+
+# the elevation under a coordinate, which is how you check a claimed camera height
+echo "4.8909 52.3738" | gdallocationinfo -wgs84 -valonly dem.tif
+```
+
`gdalinfo` is the honesty check: if it reports no CRS, the file is a picture and any measurement
you take off it is invented. Georeferencing error concentrates away from your control points, so
put them at the corners of the area you care about and expect metres of error, not centimetres.
Resampling with `-r cubic` makes imagery look better and makes pixel-level forensics worse — use
-`-r near` when the pixels themselves are the evidence.
+`-r near` when the pixels themselves are the evidence. And `gdal_calc.py` will happily difference
+two rasters of different extents unless you say `--extent=intersect`, producing a delta image whose
+bright edges are registration error rather than change.
### QGIS
@@ -201,6 +349,9 @@ qgis_process list
# the parameters for one algorithm, before you guess at them
qgis_process help native:buffer
+# machine-readable, for building a pipeline against the real parameter names
+qgis_process --json help native:extractbylocation | jq '.parameters | keys'
+
# a 200m buffer around candidate points
qgis_process run native:buffer -- INPUT=candidates.gpkg DISTANCE=200 OUTPUT=buffered.gpkg
@@ -211,14 +362,33 @@ qgis_process run native:reprojectlayer -- \
# centroids of building polygons, for matching against a geotagged photo set
qgis_process run native:centroids -- INPUT=buildings.gpkg OUTPUT=centroids.gpkg
-# keep only the features inside an area of interest
+# keep only the features inside an area of interest. PREDICATE is an enum, not a word:
+# 0 intersect, 1 contain, 2 disjoint, 3 equal, 4 touch, 5 overlap, 6 are within, 7 cross
qgis_process run native:extractbylocation -- \
INPUT=centroids.gpkg PREDICATE=0 INTERSECT=aoi.gpkg OUTPUT=inside.gpkg
+
+# units are a run-time choice, so a buffer in metres has to say so
+qgis_process run native:buffer --distance_units=meters --area_units=m2 \
+ --ellipsoid=EPSG:7030 -- INPUT=candidates.gpkg DISTANCE=200 OUTPUT=buffered.gpkg
+
+# against a project, so layer references and saved styles resolve
+qgis_process run native:centroids --project_path=case.qgz -- \
+ INPUT=buildings.gpkg OUTPUT=centroids.gpkg
+
+# parameters as JSON on stdin, which is how this goes into a script without quoting pain
+echo '{"inputs": {"INPUT": "candidates.gpkg", "DISTANCE": 200, "OUTPUT": "buffered.gpkg"}}' \
+ | qgis_process run native:buffer -
+
+# startup is dominated by plugin loading; skip it for a batch of a hundred runs
+qgis_process --no-python --skip-loading-plugins run native:centroids -- \
+ INPUT=buildings.gpkg OUTPUT=centroids.gpkg
```
Measurements in a geographic CRS (`EPSG:4326`) are in degrees, not metres, and QGIS will happily
give you a meaningless number. Reproject to a local metric CRS or a UTM zone before you measure
-anything you intend to publish, and say which CRS you used.
+anything you intend to publish, and say which CRS you used. `--skip-loading-plugins` is not free of
+consequence either: an algorithm provided by a plugin disappears from `list` when you pass it, and
+the failure reads as a missing algorithm rather than a missing plugin.
### Copernicus Data Space
@@ -257,11 +427,43 @@ curl -s -L -H "Authorization: Bearer $ACCESS_TOKEN" \
"$CAT(08f7cbba-56c6-4730-b2d1-63ee8a5b9536)/\$value" -o product.zip
```
+Three refinements turn that from a listing into a search. `contains(Name,'MSIL2A')` restricts to
+the atmospherically corrected processing level, which is the only level you may compare between
+dates. The attribute form filters on cloud cover, which is the single number that decides whether a
+scene is worth downloading. And `$count=True` reports the total so you learn how many scenes exist
+before you page through them.
+
+```bash
+POLY="POLYGON((4.885 52.370,4.897 52.370,4.897 52.378,4.885 52.378,4.885 52.370))"
+
+# L2A only, over the polygon, under 10 per cent cloud, oldest first, with a total count.
+# The attribute syntax is verbose and exact: the value type appears twice.
+curl -s -G "$CAT" \
+ --data-urlencode "\$filter=Collection/Name eq 'SENTINEL-2' and contains(Name,'MSIL2A') and OData.CSC.Intersects(area=geography'SRID=4326;$POLY') and ContentDate/Start gt 2026-05-01T00:00:00.000Z and ContentDate/Start lt 2026-06-01T00:00:00.000Z and Attributes/OData.CSC.DoubleAttribute/any(att:att/Name eq 'cloudCover' and att/OData.CSC.DoubleAttribute/Value lt 10.00)" \
+ --data-urlencode '$orderby=ContentDate/Start asc' \
+ --data-urlencode '$count=True' --data-urlencode '$top=3' \
+ | jq -r '"count: \(.["@odata.count"])", (.value[] | "\(.ContentDate.Start[0:19]) \(.Name)")'
+
+# paging: the response carries @odata.nextLink, already signed with every parameter.
+# Follow it rather than incrementing $skip by hand.
+curl -s -G "$CAT" \
+ --data-urlencode "\$filter=Collection/Name eq 'SENTINEL-1' and ContentDate/Start gt 2026-09-01T00:00:00.000Z" \
+ --data-urlencode '$count=True' --data-urlencode '$top=20' \
+ | jq -r '.["@odata.nextLink"]'
+
+# the Id, which is the only thing the download endpoint accepts
+curl -s -G "$CAT" \
+ --data-urlencode "\$filter=contains(Name,'S2A_MSIL2A_20260501T104651')" \
+ | jq -r '.value[] | "\(.Id) \(.ContentLength) \(.Online)"'
+```
+
Filter on `Collection/Name` or the query will be slow enough to time out. A Sentinel-2 product
name encodes the tile and the processing level — `L2A` is atmospherically corrected and what you
want for comparing two dates; `L1C` is not. The cloud percentage in the metadata is for the whole
-100km tile, so a "12% cloud" scene can still be solid cloud over your target. Tokens expire in
-minutes; re-request rather than caching them.
+100km tile, so a "12% cloud" scene can still be solid cloud over your target, and the inverse also
+bites: a 60 per cent scene can be perfectly clear over your 500m box. Filter on it to rank, never
+to exclude outright. `Online: false` means the product has been moved to cold storage and the
+download will stall rather than fail. Tokens expire in minutes; re-request rather than caching them.
### Google Earth Pro
@@ -284,10 +486,148 @@ Import KML/GPX drop in Overpass output or a track for overlay
Status-bar date stamp the capture date of the imagery currently drawn
```
+There is no command line, so the scriptable surface is KML, which Earth Pro reads and writes. Two
+documents are worth keeping as templates. A `Placemark` with a `LookAt` reproduces an exact view —
+position, bearing and camera distance — so a colleague opens the same frame rather than the same
+coordinate. A `GroundOverlay` drapes a georeferenced scan or a cropped Sentinel tile over the
+imagery at a stated extent, which is how you compare your own raster against Earth Pro's archive
+without leaving Earth Pro.
+
+```xml
+<?xml version="1.0" encoding="UTF-8"?>
+<kml xmlns="http://www.opengis.net/kml/2.2">
+ <Document>
+ <!-- reproduces a view, not just a point: heading is the bearing the camera
+ faces, tilt 0 is straight down, range is metres from the target -->
+ <Placemark>
+ <name>Warehouse, south elevation</name>
+ <LookAt>
+ <longitude>4.8909</longitude>
+ <latitude>52.3738</latitude>
+ <altitude>0</altitude>
+ <heading>312</heading>
+ <tilt>65</tilt>
+ <range>400</range>
+ </LookAt>
+ <TimeSpan>
+ <begin>2026-05-01</begin>
+ <end>2026-09-30</end>
+ </TimeSpan>
+ <Point><coordinates>4.8909,52.3738,0</coordinates></Point>
+ </Placemark>
+
+ <!-- drape your own raster over Earth Pro's imagery. The LatLonBox must match
+ the file's real extent: gdalinfo prints it, and guessing it shifts
+ everything you then "measure" off the overlay -->
+ <GroundOverlay>
+ <name>Sentinel-2 L2A crop, 2026-05-01</name>
+ <color>b4ffffff</color>
+ <drawOrder>1</drawOrder>
+ <Icon><href>crop_wgs84.png</href></Icon>
+ <altitudeMode>clampToGround</altitudeMode>
+ <LatLonBox>
+ <north>52.378</north>
+ <south>52.370</south>
+ <east>4.897</east>
+ <west>4.885</west>
+ <rotation>0</rotation>
+ </LatLonBox>
+ </GroundOverlay>
+ </Document>
+</kml>
+```
+
+```bash
+# produce the overlay PNG the KML above expects: WGS84, so the LatLonBox is honest
+gdalwarp -t_srs EPSG:4326 -r near crop.tif crop_wgs84.tif
+gdal_translate -of PNG crop_wgs84.tif crop_wgs84.png
+
+# read the extent back out, in the order the LatLonBox wants it
+gdalinfo -json crop_wgs84.tif \
+ | jq -r '.wgs84Extent.coordinates[0] | "west \(.[0][0]) south \(.[0][1]) east \(.[2][0]) north \(.[2][1])"'
+```
+
+Sharing a view outside Earth Pro has a documented URL form as well, which is the fastest way to
+hand someone a Street View frame at a specific bearing rather than a coordinate they then have to
+orient themselves in.
+
+```text
+https://www.google.com/maps/@?api=1&map_action=pano&viewpoint=52.3738,4.8909&heading=312&pitch=0&fov=80
+https://www.google.com/maps/@?api=1&map_action=map¢er=52.3738,4.8909&zoom=18&basemap=satellite
+
+ api=1 required, and the URL silently misbehaves without it
+ viewpoint lat,lon -- Google snaps to the nearest panorama, which may be the wrong street
+ pano a specific panorama id, when you need that exact capture and not the nearest
+ heading -180 to 360, compass bearing the camera faces
+ pitch -90 to 90, 0 is horizontal
+ fov 10 to 100, default 90 -- narrowing it is the closest thing to a zoom
+ zoom 0 to 21 on map_action=map
+ basemap roadmap, satellite or terrain
+```
+
The displayed date is the date of the *dominant* image in the view; a mosaic can blend captures
months apart, with a visible seam. Zooming changes which image is drawn, so the date can change
under you without the view appearing to move. 3D buildings are models, not imagery, and are not
-evidence of anything.
+evidence of anything. A `viewpoint` link resolves to whatever panorama is nearest at the time
+someone opens it, so for anything you intend to cite, capture the `pano` id instead — the nearest
+panorama changes when Google drives the street again.
+
+### Google Earth Engine
+
+The free analysis platform for the same archives, and the right tool when the question spans more
+scenes than you want to download. The [code editor](https://code.earthengine.google.com/) runs
+JavaScript server-side against the full Sentinel and Landsat catalogues; a free account is needed
+and non-commercial use is free. The pattern is always the same: a geometry, a collection, filters,
+then a composite or a difference.
+
+```javascript
+// an area of interest, not a point: filterBounds takes any geometry
+var aoi = ee.Geometry.Rectangle([4.885, 52.370, 4.897, 52.378]);
+
+// the harmonised collection, because the pre-2022 and post-2022 scenes are
+// otherwise offset by a processing baseline change and every difference is wrong
+var s2 = ee.ImageCollection('COPERNICUS/S2_SR_HARMONIZED')
+ .filterBounds(aoi)
+ .filterDate('2026-05-01', '2026-06-01')
+ .filter(ee.Filter.lt('CLOUDY_PIXEL_PERCENTAGE', 20));
+
+print('scenes matched', s2.size());
+
+// the least cloudy scene in the window, rather than the first one returned
+var best = ee.Image(s2.sort('CLOUDY_PIXEL_PERCENTAGE').first());
+print('chosen', best.get('system:index'), best.get('CLOUDY_PIXEL_PERCENTAGE'));
+
+Map.setCenter(4.891, 52.374, 16);
+Map.addLayer(best, {bands: ['B4', 'B3', 'B2'], min: 0, max: 3000}, 'true colour');
+
+// false colour: vegetation bright red, bare ground and new works pale.
+// This is the composite that makes a demolition or an earthwork obvious at 10m.
+Map.addLayer(best, {bands: ['B8', 'B4', 'B3'], min: 0, max: 4000}, 'false colour');
+
+// a two-date NDVI difference, which is change detection without downloading anything
+var ndvi = function (img) {
+ return img.normalizedDifference(['B8', 'B4']).rename('ndvi');
+};
+var may = ndvi(ee.Image(s2.sort('CLOUDY_PIXEL_PERCENTAGE').first()));
+var sep = ndvi(ee.Image(
+ ee.ImageCollection('COPERNICUS/S2_SR_HARMONIZED')
+ .filterBounds(aoi)
+ .filterDate('2026-09-01', '2026-10-01')
+ .filter(ee.Filter.lt('CLOUDY_PIXEL_PERCENTAGE', 40))
+ .sort('CLOUDY_PIXEL_PERCENTAGE')
+ .first()));
+
+Map.addLayer(sep.subtract(may).clip(aoi),
+ {min: -0.6, max: 0.6, palette: ['red', 'white', 'green']}, 'NDVI delta');
+```
+
+The thing to distrust here is the compositing. A median over a date range looks clean and is not an
+image of any moment — it is a statistic, and it cannot carry a capture date, so it supports no
+time-sensitive claim. Pick a single scene and name it when the claim is about a date.
+`CLOUDY_PIXEL_PERCENTAGE` is the whole-tile figure again, so a 15 per cent scene can still be
+clouded over your box; look at the picture before you trust the number. And `Map.addLayer` stretches
+reflectance to a `min`/`max` you chose, so two dates displayed with different stretches will look
+different whether or not anything changed.
### Mapillary
@@ -300,7 +640,7 @@ makes "what did this junction look like in March" a scriptable question.
TOKEN='MLY|xxxx|xxxx'
API='https://graph.mapillary.com'
-# images in a bounding box (minLon,minLat,maxLon,maxLat) — must be under 0.01 degrees square
+# images in a bounding box (left,bottom,right,top) — must be under 0.01 degrees square
curl -s -H "Authorization: OAuth $TOKEN" \
"$API/images?bbox=4.890,52.372,4.895,52.375&fields=id,captured_at,compass_angle,geometry&limit=50" \
| jq -r '.data[] | "\(.captured_at) \(.id)"'
@@ -322,13 +662,47 @@ curl -s -H "Authorization: OAuth $TOKEN" "$API/IMAGE_ID?fields=thumb_2048_url" \
| jq -r .thumb_2048_url | xargs curl -s -o frame.jpg
```
+Four more parameters are worth knowing. `is_pano` separates 360-degree captures, which are the
+ones where you can look behind the camera; `creator_username` and `organization_id` let you decide
+how much to trust a sequence by who uploaded it; and `sequence_ids` plus the `image_ids` endpoint
+walks a single drive in capture order, which is what turns a set of frames into a direction of
+travel.
+
+```bash
+# 360 captures only, with the computed fields alongside the raw ones
+curl -s -H "Authorization: OAuth $TOKEN" \
+ "$API/images?bbox=4.890,52.372,4.895,52.375&is_pano=true&fields=id,captured_at,is_pano,compass_angle,computed_compass_angle,computed_geometry&limit=100" \
+ | jq -r '.data[] | "\(.captured_at) raw \(.compass_angle) computed \(.computed_compass_angle)"'
+
+# who uploaded what here: trust a municipal fleet differently from one contributor
+curl -s -H "Authorization: OAuth $TOKEN" \
+ "$API/images?bbox=4.890,52.372,4.895,52.375&fields=id,captured_at,creator,organization" \
+ | jq -r '.data[] | "\(.captured_at) \(.creator.username // "-") \(.organization // "-")"'
+
+# every image id in one sequence, in capture order. This endpoint ignores `fields`.
+curl -s -H "Authorization: OAuth $TOKEN" "$API/image_ids?sequence_id=SEQUENCE_ID" \
+ | jq -r '.data[].id' | head -20
+
+# the sequence a frame belongs to, then the frames either side of it
+curl -s -H "Authorization: OAuth $TOKEN" "$API/IMAGE_ID?fields=sequence,captured_at"
+
+# detections on one frame as a separate edge, with geometry per detection
+curl -s -H "Authorization: OAuth $TOKEN" \
+ "$API/IMAGE_ID/detections?fields=value,geometry,created_at" \
+ | jq -r '.data[] | .value' | sort | uniq -c | sort -rn
+```
+
The bounding box limit is real: anything larger than 0.01 degrees square is rejected rather than
-truncated, so tile your area. `compass_angle` is the camera bearing and is what lets you say which
-side of the street a feature is on. Coverage is contributor-driven, so a road can have 2019 and
-2026 imagery and nothing between, and sequence positions are GPS traces — metres of error in
-cities, more in canyons.
+truncated, so tile your area. `limit` defaults to 2000 and tops out there, so a dense city tile can
+be truncated without an error — compare the count you get against the count you expected rather
+than assuming you have everything. `compass_angle` is the camera bearing and is what lets you say
+which side of the street a feature is on; `computed_compass_angle` and `computed_geometry` are the
+structure-from-motion refinements and are usually the better of the two, but they exist only for
+frames that were successfully reconstructed, so a `null` there is a frame whose position is raw GPS.
+Coverage is contributor-driven, so a road can have 2019 and 2026 imagery and nothing between, and
+sequence positions are GPS traces — metres of error in cities, more in canyons.
-### EarthExplorer
+### EarthExplorer and the LandsatLook STAC API
Web only, free with registration, and the only practical route to two archives nothing else
carries: Landsat back to 1972, and declassified US reconnaissance imagery from the 1960s to the
@@ -345,11 +719,162 @@ and the order matters:
4 Results the footprint icon shows coverage; the browse icon previews it
```
-Register before you search, because the download buttons are hidden until you log in, and read
-the scene's entity ID — it encodes the mission and date, and it is what you cite. Declassified
-frames arrive as scanned film with no georeferencing at all, which is where the
+For the Landsat half of that, there is a scriptable route that needs no account at all. USGS runs a
+STAC API at `landsatlook.usgs.gov/stac-server`, and search is open. This is the fastest way to find
+out whether a usable Landsat scene exists over a point in a window, and the response carries the
+sun azimuth and elevation per scene, which is the shadow geometry you need for
+[chronolocation](/sheets/osint/geolocation) without computing anything.
+
+```bash
+STAC='https://landsatlook.usgs.gov/stac-server'
+
+# what collections exist, and their ids. landsat-c2l2-sr is Level-2 surface reflectance.
+curl -s "$STAC/collections" | jq -r '.collections[] | "\(.id) \(.title)"'
+
+# a search: bbox, date window, cloud ceiling, oldest first
+curl -s -X POST "$STAC/search" -H 'Content-Type: application/json' -d '{
+ "collections": ["landsat-c2l2-sr"],
+ "bbox": [4.885, 52.370, 4.897, 52.378],
+ "datetime": "2026-05-01T00:00:00Z/2026-09-30T23:59:59Z",
+ "query": {"eo:cloud_cover": {"lt": 20}},
+ "sortby": [{"field": "properties.datetime", "direction": "asc"}],
+ "limit": 5
+}' | jq -r '"matched: \(.numberMatched)",
+ (.features[] | "\(.properties.datetime[0:19]) \(.id) cloud \(.properties["eo:cloud_cover"]) sun_az \(.properties["view:sun_azimuth"]) sun_el \(.properties["view:sun_elevation"])")'
+
+# the asset list for one scene: per-band COGs, addressable without downloading the scene
+curl -s "$STAC/collections/landsat-c2l2-sr/items/LC09_L2SP_199023_20260528_20260530_02_T1_SR" \
+ | jq -r '.assets | to_entries[] | "\(.key) \(.value.href)"' | head -20
+```
+
+The same API from a shell, via `pystac-client`, which handles paging and saves a searchable item
+collection. `--matched` asks only for the count, which is the cheap question to ask first.
+
+```bash
+pipx install pystac-client # or: pip install pystac-client
+
+# how many scenes exist, before fetching any of them
+stac-client search "$STAC" -c landsat-c2l2-sr \
+ --bbox 4.885 52.370 4.897 52.378 --datetime 2026-05-01/2026-09-30 --matched
+
+# the same search, cloud-filtered, newest first, saved to a file
+stac-client search "$STAC" -c landsat-c2l2-sr \
+ --bbox 4.885 52.370 4.897 52.378 --datetime 2026-05-01/2026-09-30 \
+ --query "eo:cloud_cover<20" --sortby "-properties.datetime" \
+ --max-items 20 --save landsat_items.json
+
+# only the fields you need, piped straight into a table
+stac-client search "$STAC" -c landsat-c2l2-sr \
+ --bbox 4.885 52.370 4.897 52.378 --datetime 2026-07-01/2026-07-31 \
+ --fields "id,properties.datetime,properties.eo:cloud_cover" \
+ | jq -r '.features[] | [.id, .properties.datetime, .properties["eo:cloud_cover"]] | @tsv'
+
+# an arbitrary polygon instead of a box: a GeoJSON file or an inline geometry
+stac-client search "$STAC" -c landsat-c2l2-sr --intersects aoi.geojson \
+ --datetime 2026-01-01/2026-12-31 --matched
+```
+
+Register before you search EarthExplorer, because the download buttons are hidden until you log in,
+and read the scene's entity ID — it encodes the mission and date, and it is what you cite.
+Declassified frames arrive as scanned film with no georeferencing at all, which is where the
`gdal_translate -gcp` workflow above earns its keep. Landsat Collection 2 Level-2 is corrected and
-comparable between dates; Level-1 is not, so do not difference the two.
+comparable between dates; Level-1 is not, so do not difference the two. The STAC route covers only
+Landsat: the declassified archive is not in it, so that half of the job stays in the web interface.
+
+### NASA FIRMS
+
+Active-fire detections from MODIS and VIIRS, as a CSV you can query by box and date. This is the
+only source on this page that answers "was there a fire or an explosion here, and when" to the
+hour, because the thermal sensors pass several times a day where the optical archive manages one
+cloud-free frame a week. A free map key from the FIRMS site is the only requirement.
+
+```bash
+KEY='your_map_key'
+FIRMS='https://firms.modaps.eosdis.nasa.gov/api'
+
+# which dates each product actually covers, before you ask for one that does not exist
+curl -s "$FIRMS/data_availability/csv/$KEY/ALL"
+
+# detections in an area over the last 3 days. AREA is west,south,east,north --
+# the opposite order to Overpass, and the usual source of an empty result set
+curl -s "$FIRMS/area/csv/$KEY/VIIRS_SNPP_NRT/34.0,31.0,36.5,33.5/3"
+
+# a historical window: same path plus a start date, YYYY-MM-DD. Day range is 1 to 5,
+# so a month is six requests, not one.
+curl -s "$FIRMS/area/csv/$KEY/VIIRS_NOAA20_NRT/34.0,31.0,36.5,33.5/5/2026-09-01"
+
+# the whole globe for a day, when you do not yet know where to look
+curl -s "$FIRMS/area/csv/$KEY/MODIS_NRT/world/1"
+
+# high-confidence night-time detections only, sorted by brightness
+curl -s "$FIRMS/area/csv/$KEY/VIIRS_SNPP_NRT/34.0,31.0,36.5,33.5/3" \
+ | awk -F, 'NR==1 || ($14=="h" && $12=="N")' | sort -t, -k3 -rn | head
+```
+
+Read the columns before you read the map. `confidence` is `l`/`n`/`h` for VIIRS and a percentage
+for MODIS, `daynight` is `D` or `N`, `frp` is fire radiative power in megawatts, and `scan`/`track`
+give the pixel footprint — which is 375m for VIIRS and 1km for MODIS, so a detection is an area, not
+a point, and plotting it as a dot overstates your precision by hundreds of metres. The `_NRT`
+sources are near-real-time and get reprocessed; the `_SP` standard-product versions are the ones to
+cite weeks later, and they will not agree exactly. A detection is a thermal anomaly: gas flares,
+industrial furnaces and sunglint off metal roofs all produce them, so the question a FIRMS hit
+answers is "was something hot here at 22:14 UTC", not "was there an airstrike". Conflict-specific
+use of this feed is on [Conflict & Environment](/sheets/osint/conflict-and-environment).
+
+### NASA Worldview and GIBS
+
+Worldview is the browser ([worldview.earthdata.nasa.gov](https://worldview.earthdata.nasa.gov/))
+and GIBS is the tile service behind it. The resolution is coarse — 250m at best — and that is not
+the point: GIBS serves a dated, global, cloud-free-ish image for *every single day* back years,
+with no key and no account, which no other source on this page does. It is how you establish what
+the weather was doing on the day your high-resolution scene is missing.
+
+```bash
+GIBS='https://gibs.earthdata.nasa.gov'
+
+# one WMTS tile, REST form:
+# /wmts/{projection}/best/{layer}/default/{time}/{tilematrixset}/{z}/{row}/{col}.{ext}
+curl -s -o tile.jpg \
+ "$GIBS/wmts/epsg4326/best/VIIRS_SNPP_CorrectedReflectance_TrueColor/default/2026-09-20/250m/6/13/36.jpg"
+
+# a bounded image via WMS instead, which is what you want for an area of interest.
+# Version 1.3.0 uses CRS and, for EPSG:4326, BBOX in lat,lon order: south,west,north,east
+curl -s -o scene.png "$GIBS/wms/epsg4326/best/wms.cgi?\
+version=1.3.0&service=WMS&request=GetMap&format=image/png&STYLE=default\
+&CRS=EPSG:4326&BBOX=52.3,4.8,52.5,5.0&WIDTH=1200&HEIGHT=1200&TIME=2026-09-20\
+&LAYERS=VIIRS_SNPP_CorrectedReflectance_TrueColor"
+
+# the layer catalogue, including each layer's available date range
+curl -s "$GIBS/wmts/epsg4326/best/1.0.0/WMTSCapabilities.xml" \
+ | grep -oE '<ows:Identifier>[^<]+' | sed 's/.*>//' | head -40
+```
+
+The WMS endpoint is a GDAL data source, so the day-by-day archive can be pulled straight into the
+same pipeline as everything else, which beats screenshotting the browser:
+
+```bash
+# GDAL reads the WMS URL directly, so crop and reproject in one step
+gdal_translate -of GTiff -projwin 4.8 52.5 5.0 52.3 -projwin_srs EPSG:4326 \
+ "WMS:$GIBS/wms/epsg4326/best/wms.cgi?LAYERS=VIIRS_SNPP_CorrectedReflectance_TrueColor&TIME=2026-09-20&SRS=EPSG:4326&FORMAT=image/png&VERSION=1.1.1" \
+ day.tif
+
+# a week of daily frames, to find the cloud-free day worth paying attention to
+for d in 2026-09-{14..20}; do
+ curl -s -o "gibs_$d.png" "$GIBS/wms/epsg4326/best/wms.cgi?\
+version=1.3.0&service=WMS&request=GetMap&format=image/png&STYLE=default\
+&CRS=EPSG:4326&BBOX=52.0,4.0,53.0,5.5&WIDTH=800&HEIGHT=800&TIME=$d\
+&LAYERS=VIIRS_SNPP_CorrectedReflectance_TrueColor"
+done
+```
+
+Projections are separate endpoints — `epsg4326`, `epsg3857`, `epsg3413` and `epsg3031` — and a
+layer present in one is not necessarily present in another, with the polar projections carrying far
+fewer. WMS 1.3.0 reverses the `BBOX` axis order for EPSG:4326 relative to 1.1.1, which is the
+reason a request returns ocean when you asked for land; if the image looks like the wrong
+hemisphere, swap the pairs rather than doubting the coordinates. `TIME` is a date, not a timestamp,
+and the frame you get is the composite for that day's overpasses, so a "2026-09-20" image spans
+hours. At 250m a building is a fifth of a pixel: use this to date weather and smoke plumes, never
+to identify a structure.
### SunCalc and ShadeMap
@@ -371,10 +896,26 @@ Capture: the coordinate, date, computed azimuth and elevation, and the tool's ow
Caveat: a date you have not independently established makes the whole result circular
```
+SunCalc keeps its entire state in the URL fragment, which is the only part of it worth treating as
+a command line. The form is `#/<lat>,<lon>,<zoom>/<YYYY.MM.DD>/<HH:MM>/<object height>`, so a link
+reproduces a specific reading rather than just opening the tool:
+
+```text
+https://www.suncalc.org/#/52.3738,4.8909,17/2026.07.14/17:40/5/2
+
+ 52.3738,4.8909 the coordinate
+ 17 map zoom
+ 2026.07.14 date, dot-separated
+ 17:40 local time at that coordinate, in the timezone the page reports
+ 5 object height in metres, which drives the shadow-length readout
+```
+
Both assume you have the location right; a 50m error in position barely moves the azimuth, but a
wrong date moves it by degrees per week near the solstices. Shadow work narrows a time, it does
not prove one — pair it with [image and video forensics](/sheets/osint/image-video-forensics) and
-with whatever the metadata claims.
+with whatever the metadata claims. The full arithmetic, the two-date-window problem and the Python
+libraries that sweep a whole year are on
+[Geolocation & Chronolocation](/sheets/osint/geolocation).
## Tool reference
@@ -445,10 +986,36 @@ with whatever the metadata claims.
- **Undated imagery is useless for a time-sensitive claim.** Record the capture date every time.
- **Cloud cover ruins optical revisit rates.** A 5-day nominal revisit can mean a month of usable
imagery in the wet season. Radar (Sentinel-1) sees through cloud but is much harder to read.
+- **A scene's cloud percentage is for the whole tile, not your target.** It is a 100km tile for
+ Sentinel-2 and a 185km swath for Landsat. Use the number to rank candidates and then look at the
+ picture; a 60 per cent scene can be clear over your box and a 12 per cent scene can be solid
+ cloud over it.
- **OSM is crowd-sourced.** Completeness varies enormously by region, and an absent feature may
simply be unmapped.
- **Basemap labels disagree**, particularly on disputed borders and place names. Say which source
you used.
+- **Bounding boxes are in four different orders across these tools.** Overpass takes
+ `south,west,north,east`; Mapillary and STAC take `left,bottom,right,top`; FIRMS takes
+ `west,south,east,north`; and WMS 1.3.0 flips to lat,lon for EPSG:4326 where 1.1.1 does not. An
+ empty result set is this mistake far more often than it is an absence of data.
+- **A median composite has no capture date.** It is a statistic over a window, so it cannot support
+ a claim about a day. Pick a single scene and name it.
+- **Processing levels are not comparable.** L2A against L1C, or Landsat Level-2 against Level-1,
+ produces a difference image of the atmospheric correction rather than of the ground.
+- **Differencing two rasters of different extents aligns them silently.** `gdal_calc.py` without
+ `--extent=intersect` will give you a delta whose brightest features are registration error.
+- **A 10m pixel cannot resolve a 10m object.** Two or three pixels across is the floor for seeing
+ that something is there; you need an order of magnitude better to say what it is. A roof six
+ pixels wide is enough to see it disappear and not enough to see how.
+- **Street-level coverage is a sample, not a survey.** A junction with 2019 and 2026 frames and
+ nothing between does not mean nothing happened in between, and the newest frame is not the frame
+ nearest your date.
+- **A thermal detection is an area, not a point.** VIIRS pixels are 375m and MODIS 1km, so plotting
+ a FIRMS hit as a dot claims a precision the sensor does not have.
+- **Near-real-time products get reprocessed.** A FIRMS `_NRT` detection and the later `_SP` version
+ of the same pass will not agree exactly, so cite the one you can still retrieve.
+- **Google's historical slider date is the dominant image's date.** A mosaic blends captures months
+ apart, and changing zoom can change which image is drawn without the view appearing to move.
## Worked example
@@ -457,30 +1024,91 @@ that a warehouse there was demolished in the summer of 2026.
1. **Establish what is mapped.** An Overpass query for `building` ways in a small box around the
coordinate returns three polygons, one tagged `building=warehouse` with an `addr:street`. That
- gives the feature a name and an address to search on.
+ gives the feature a name and an address to search on. The `(newer:"2026-06-01T00:00:00Z")`
+ form on the same query shows one of the three edited in August 2026, which is a free first
+ corroboration of the claim's timing from OSM's own history.
2. **Find free imagery either side of the claim.** The Copernicus OData catalogue, filtered to
- `SENTINEL-2` and intersected with a polygon around the point, lists L2A products for late May
- and early September 2026. Two dates bracketing the claim is the minimum useful set.
+ `SENTINEL-2`, `contains(Name,'MSIL2A')` and intersected with a 1.3km polygon around the point.
+ Run once per month with a 10 per cent cloud ceiling, and once without, because the gap between
+ those two answers is the real story:
+
+```text
+=== MAY 2026, cloudCover < 10 ===
+count: 3
+ 2026-05-01T10:36:19 S2B_MSIL2A_20260501T103619_N0512_R008_T31UFU_20260501T143617.SAFE
+ 2026-05-01T10:46:51 S2A_MSIL2A_20260501T104651_N0512_R051_T31UFU_20260501T173800.SAFE
+ 2026-05-26T10:36:21 S2C_MSIL2A_20260526T103621_N0512_R008_T31UFU_20260526T140311.SAFE
+
+=== SEPTEMBER 2026, cloudCover < 10 ===
+count: 0
+
+=== SEPTEMBER 2026, cloudCover < 40 ===
+count: 2
+ 2026-09-01T10:46:19 S2B_MSIL2A_20260901T104619_N0512_R051_T31UFU_20260901T131914.SAFE
+ 2026-09-25T10:40:41 S2A_MSIL2A_20260925T104041_N0513_R008_T31UFU_20260925T171206.SAFE
+
+=== SEPTEMBER 2026, no cloud filter ===
+count: 18
+```
+
+ Eighteen scenes in the month, two under 40 per cent cloud, none under 10. The nominal five-day
+ revisit is an orbital fact; the usable cadence over the Netherlands in September is a fortnight.
+ Take the 25 September scene and accept that it needs looking at rather than trusting.
3. **Crop and compare.** `gdal_translate -projwin` cuts both scenes to the same 500m box,
`gdalwarp -t_srs EPSG:3857` puts them in the same CRS, and the pair opened in QGIS shows the
- roof present in May and bare ground in September. At 10m the roof is six pixels across — enough
- to see it go, not enough to see how.
-4. **Confirm at resolution.** Google Earth Pro's historical slider over the same point has a
- July 2026 frame at sub-metre scale showing partial demolition and plant on site. Record the
- status-bar date, not the date you looked.
-5. **Check the ground.** Mapillary `images?bbox=…&start_captured_at=2026-08-01T00:00:00Z` returns
- a contributor sequence from August with `compass_angle` facing the plot: hoarding up, structure
- gone. Street level dates the end of the work more precisely than any satellite pass.
-6. **Measure, then say so.** Reprojected to UTM, the QGIS measure tool puts the cleared footprint
- at 1,840 m², consistent with the OSM polygon's area. Quote the CRS alongside the number.
-7. **Time the photograph, if it matters.** The caption claims mid-July. SunCalc for the coordinate
- on 14 July returns an azimuth matching the shadow bearing in the image at around 17:40 local —
- consistent, not proof, and recorded as such.
-
-What you can assert: a structure present on a dated 10m scene in May and absent in September, with
-a sub-metre frame in July showing demolition in progress and street-level imagery in August
-showing it finished. What you cannot: who did it, or why — no imagery source on this page
-carries that.
+ roof present on 1 May and bare ground on 25 September. At 10m the roof is six pixels across —
+ enough to see it go, not enough to see how. `gdallocationinfo -wgs84 -valonly` on the NDVI
+ difference at the coordinate returns **+0.02**, confirming what the eye says: this is roof
+ giving way to bare ground, not vegetation change.
+4. **Bracket it with Landsat, for the sun geometry.** The LandsatLook STAC search over the same box
+ needs no account and returns, for May to September with a 20 per cent cloud ceiling:
+
+```text
+matched: 12
+2026-05-28T10:38:56 LC09_L2SP_199023_20260528_20260530_02_T1_SR cloud 4.51 sun_az 153.53 sun_el 56.26
+2026-06-22T10:33:16 LC09_L2SP_198024_20260622_20260623_02_T1_SR cloud 7.52 sun_az 148.59 sun_el 58.83
+2026-06-29T10:39:07 LC09_L2SP_199023_20260629_20260630_02_T1_SR cloud 6.09 sun_az 150.22 sun_el 57.48
+2026-07-15T10:39:14 LC09_L2SP_199023_20260715_20260717_02_T1_SR cloud 0.22 sun_az 150.28 sun_el 55.66
+```
+
+ 30m is too coarse to see the building, so this is not the change-detection source here. What it
+ gives you is `view:sun_elevation` around **55.7 degrees** at 10:39 UTC in mid-July, which is the
+ number to check the photograph's shadows against in step 7.
+5. **Confirm at resolution.** Google Earth Pro's historical slider over the same point has a
+ **July 2026** frame at sub-metre scale showing partial demolition and plant on site. Record the
+ status-bar date, not the date you looked, and capture the view as a `LookAt` placemark so the
+ frame is reproducible rather than described.
+6. **Check the ground.** Mapillary `images?bbox=4.8900,52.3730,4.8915,52.3745&start_captured_at=2026-08-01T00:00:00Z`
+ returns a contributor sequence from August with `computed_compass_angle` 312 degrees, facing the
+ plot: hoarding up, structure gone. Street level dates the end of the work more precisely than any
+ satellite pass. The sequence runs north-west along the street, so the frames either side
+ establish that the hoarding is on the plot boundary and not on the one next door.
+7. **Measure, then say so.** Reprojected to UTM zone 31N (`EPSG:32631`), the QGIS measure tool puts
+ the cleared footprint at **1,840 m²**, against **1,795 m²** for the OSM polygon — a 2.5 per cent
+ disagreement, which is within what a 10m pixel edge can produce and is therefore consistent
+ rather than confirming. Quote the CRS alongside the number.
+8. **Time the photograph, if it matters.** The caption claims mid-July. SunCalc for the coordinate
+ on 14 July, with the lamp standard's 5m height entered, returns an azimuth matching the shadow
+ bearing at around **17:40 local** and a solar elevation of about 30 degrees — consistent with
+ the Landsat-derived geometry for the same week, and recorded as consistent rather than proven.
+
+What you can assert: a structure present on a named, dated 10m scene on 1 May 2026 and absent from
+a named, dated 10m scene on 25 September 2026; a sub-metre Google Earth Pro frame stamped July 2026
+showing demolition in progress; a Mapillary sequence from August 2026 showing the site hoarded and
+cleared; and a cleared footprint of 1,840 m² in EPSG:32631. The demolition therefore falls between
+1 May and 25 September, and the July frame narrows it to the first half of that window. What you
+cannot assert: who did it, or why — no imagery source on this page carries that.
+
+What would falsify it: a July Google Earth Pro frame that turns out to be a mosaic blending a
+pre-demolition capture from a neighbouring strip, which the visible seam would show and the
+status-bar date would not; a Mapillary sequence whose `computed_geometry` is absent, leaving the
+position as raw GPS and the "facing the plot" claim unsupported; or a warehouse that was rebuilt
+and re-demolished, which two dated frames five months apart cannot distinguish from one event. The
+measurement carrying the most risk is the **1,840 m² footprint**. It is traced off 10m pixels, so
+each edge carries at best half a pixel of uncertainty: on a roughly 43m square that is about ±5m
+per side, or **±8 per cent on the area**. Quote it as 1,840 m² ±150 m² or do not quote a figure at
+all — and note that this tolerance is why the 2.5 per cent agreement with the OSM polygon in step 7
+corroborates nothing. Two numbers that agree inside their error bars are not a cross-check.
## Broader catalogues
diff --git a/src/content/sheets/osint/osint-foundations.md b/src/content/sheets/osint/osint-foundations.md
@@ -649,12 +649,20 @@ exposure posture, a continuous capture log, two registry records with hashes and
explicitly rejected lookalike company, and a negative-results entry that says what "two hits" does
and does not mean.
+What you can assert: which sources were queried, from which exposure posture, what they returned,
+and that the preserved registry artefacts still match the hashes recorded at acquisition. This is
+an auditable collection record, not a finding that the tipster's allegation is true.
+
What none of that establishes: whether Northgate is a front. The tipster's claim is still exactly
one uncorroborated assertion, recorded as such, in a section of the note that cannot be mistaken
for a finding. The substantive work now moves to
[Company & Financial Records](/sheets/osint/companies-and-finance) — but it moves there on top of a
record that will survive someone attacking it, which is the only thing this sheet is for.
+What would falsify it: an exposure check showing traffic left by the host rather than the research
+VM, a gap in the action log, or a hash mismatch on either source artefact. Those failures do not
+prove the allegation false; they make the collection record too weak to support later claims.
+
## Broader catalogues
- [Foundational OSINT Tools](https://tools.osintnewsletter.com/tool-categories/foundational-osint-tools)
diff --git a/src/content/sheets/osint/people-search.md b/src/content/sheets/osint/people-search.md
@@ -4,9 +4,9 @@ description: "Registries, court records, aggregators and breach data for identif
category: osint
subcategory: "People & Identity"
tags: [osint, people, public-records, breach-data]
-tools: [hibp, intelx, aleph, courtlistener, opensanctions, ratsit]
+tools: [hibp, intelx, aleph, alephclient, courtlistener, opensanctions, yente, edgar, ratsit, hitta]
difficulty: intermediate
-updated: 2026-09-28
+updated: 2026-10-04
references:
- name: "Bellingcat's Online Investigation Toolkit"
url: "https://bellingcat.gitbook.io/toolkit"
@@ -20,6 +20,36 @@ references:
license: none
relation: derived
note: "Second tool catalogue, cross-checked against the above."
+ - name: "Have I Been Pwned API v3"
+ url: "https://haveibeenpwned.com/API/v3"
+ author: "Troy Hunt"
+ relation: link-only
+ note: "Endpoint paths, required headers and rate-limit behaviour."
+ - name: "Intelligence X SDK (Python)"
+ url: "https://github.com/IntelligenceX/SDK/tree/master/Python"
+ author: "Kleissner Investments"
+ relation: link-only
+ note: "CLI flags, API endpoint paths and the sort/media enumerations."
+ - name: "alephclient documentation"
+ url: "https://docs.aleph.occrp.org/developers/alephclient/"
+ author: "OCCRP"
+ relation: link-only
+ note: "alephclient subcommands and the ALEPHCLIENT_* environment variables."
+ - name: "CourtListener REST API v4"
+ url: "https://wiki.free.law/c/courtlistener/help/api/rest/v4/overview"
+ author: "Free Law Project"
+ relation: link-only
+ note: "Endpoint paths, search type values, throttles and the OPTIONS convention."
+ - name: "OpenSanctions API reference"
+ url: "https://api.opensanctions.org/openapi.json"
+ author: "OpenSanctions"
+ relation: link-only
+ note: "Scope paths and query parameters for search, match and statements."
+ - name: "yente"
+ url: "https://github.com/opensanctions/yente"
+ author: "OpenSanctions"
+ relation: link-only
+ note: "Self-hosted matching engine: image, port and YENTE_* environment variables."
---
## What this covers
@@ -57,6 +87,22 @@ unusually productive:
[Search Systems](https://www.searchsystems.net/) is a directory of the underlying databases,
organised by state and record type.
+None of these publish an API. They do take their search term in the query string, which is worth
+knowing because it means a registry lookup can be scripted into a browser or a note template
+rather than retyped:
+
+```text
+https://www.hitta.se/sök?vad=Anna+Andersson name or address; returns people and companies
+https://www.hitta.se/sök?vad=08-555+012+34 reverse telephone, same parameter
+```
+
+The `vad` parameter is the only one you need; the rest of the query string the site adds is UI
+state. Note that the path segment is the Swedish word `sök`, so it percent-encodes to
+`s%C3%B6k` when you paste it into `curl` — a shell that mangles the UTF-8 is the usual reason a
+hand-built Hitta URL 404s. Ratsit and the US brokers sit behind bot protection that returns 403
+to anything without a browser fingerprint, so for those, drive the search box rather than the
+URL.
+
## Key tools
### Have I Been Pwned
@@ -66,18 +112,58 @@ stable documented API. Account lookups sit behind a paid key now — the cheapes
dollars a month — while the breach catalogue and the password range endpoint stay free and need
no key at all.
+The keyless half, which is where to start because it costs nothing and tells you what the corpus
+even is:
+
```bash
# the breach catalogue: free, no key, useful for dating a corpus
-curl -s 'https://haveibeenpwned.com/api/v3/breaches' | jq -r '.[].Name' | head
+curl -s -H 'user-agent: osint-research' 'https://haveibeenpwned.com/api/v3/breaches' \
+ | jq -r '.[].Name' | head
# one breach in detail — when it happened, how big, what fields leaked
-curl -s 'https://haveibeenpwned.com/api/v3/breach/Adobe' | jq '{BreachDate,PwnCount,DataClasses}'
+curl -s -H 'user-agent: osint-research' 'https://haveibeenpwned.com/api/v3/breach/Adobe' \
+ | jq '{BreachDate,AddedDate,PwnCount,IsVerified,DataClasses}'
+
+# which breaches carried the field you actually care about, sorted oldest first
+curl -s -H 'user-agent: osint-research' 'https://haveibeenpwned.com/api/v3/breaches' \
+ | jq -r '[.[] | select(.DataClasses | index("Physical addresses"))]
+ | sort_by(.BreachDate)[] | "\(.BreachDate) \(.PwnCount) \(.Name)"'
+
+# every breach attributed to one domain — the pivot when you know the employer, not the person
+curl -s -H 'user-agent: osint-research' 'https://haveibeenpwned.com/api/v3/breaches?Domain=adobe.com'
+
+# the field vocabulary, so a report says "Physical addresses" and not "address data"
+curl -s -H 'user-agent: osint-research' 'https://haveibeenpwned.com/api/v3/dataClasses' | jq -r '.[]'
+
+# what was added most recently, for deciding whether a re-run is worth the quota
+curl -s -H 'user-agent: osint-research' 'https://haveibeenpwned.com/api/v3/latestBreach' \
+ | jq '{Name,AddedDate}'
+```
+
+Then the keyed half. `user-agent` is mandatory on every call and `hibp-api-key` on everything
+that names an account:
+
+```bash
+# confirm which tier the key actually buys before you plan a run around it
+curl -s 'https://haveibeenpwned.com/api/v3/subscription/status' \
+ -H "hibp-api-key: $HIBP_KEY" -H 'user-agent: osint-research' \
+ | jq '{SubscriptionName,Rpm,DomainSearchMaxBreachedAccounts}'
-# which breaches hold this address; both headers are mandatory
+# which breaches hold this address; truncateResponse=false is what gets you the dates
curl -s 'https://haveibeenpwned.com/api/v3/breachedAccount/target@example.com?truncateResponse=false' \
-H "hibp-api-key: $HIBP_KEY" -H 'user-agent: osint-research' \
| jq -r '.[] | "\(.BreachDate) \(.Name)"'
+# include unverified corpora — more hits, weaker provenance, so record the flag you used
+curl -s 'https://haveibeenpwned.com/api/v3/breachedAccount/target@example.com?truncateResponse=false&IncludeUnverified=true' \
+ -H "hibp-api-key: $HIBP_KEY" -H 'user-agent: osint-research' | jq 'length'
+
+# k-anonymity account search: the address never leaves your machine, only six hex characters do
+H=$(printf 'target@example.com' | shasum -a 1 | cut -c1-40 | tr 'a-f' 'A-F')
+curl -s "https://haveibeenpwned.com/api/v3/breachedaccount/range/${H:0:6}" \
+ -H "hibp-api-key: $HIBP_KEY" -H 'user-agent: osint-research' \
+ | jq -r --arg suffix "${H:6}" '.[] | select(.hashSuffix == $suffix) | .websites[]'
+
# paste sites that carried the address, which often predate the breach being named
curl -s 'https://haveibeenpwned.com/api/v3/pasteAccount/target@example.com' \
-H "hibp-api-key: $HIBP_KEY" -H 'user-agent: osint-research'
@@ -86,19 +172,53 @@ curl -s 'https://haveibeenpwned.com/api/v3/pasteAccount/target@example.com' \
curl -s 'https://haveibeenpwned.com/api/v3/breachedDomain/example.com' \
-H "hibp-api-key: $HIBP_KEY" -H 'user-agent: osint-research'
-# stealer-log hits — infostealer output, so far more recent than breach corpora (Pro tier)
+# which domains you have actually verified, when a run comes back suspiciously empty
+curl -s 'https://haveibeenpwned.com/api/v3/subscribedDomains' \
+ -H "hibp-api-key: $HIBP_KEY" -H 'user-agent: osint-research' | jq -r '.[].DomainName'
+
+# stealer-log hits — infostealer output, so far more recent than breach corpora
curl -s 'https://haveibeenpwned.com/api/v3/stealerLogsByEmail/target@example.com' \
-H "hibp-api-key: $HIBP_KEY" -H 'user-agent: osint-research'
-# check a password without sending it: only the first five SHA-1 characters leave your machine
-printf 'hunter2' | shasum | tr 'a-f' 'A-F' | cut -c1-5 \
+# the same, across a verified domain: which of your subject's colleagues ran malware
+curl -s 'https://haveibeenpwned.com/api/v3/stealerLogsByEmailDomain/example.com' \
+ -H "hibp-api-key: $HIBP_KEY" -H 'user-agent: osint-research'
+
+# which sites a stealer-infected machine had credentials for — a browsing-history proxy
+curl -s 'https://haveibeenpwned.com/api/v3/stealerLogsByWebsiteDomain/example.com' \
+ -H "hibp-api-key: $HIBP_KEY" -H 'user-agent: osint-research'
+```
+
+The password range endpoint needs no key, is not rate limited, and never sees the password:
+
+```bash
+# only the first five SHA-1 characters leave your machine
+printf 'hunter2' | shasum -a 1 | tr 'a-f' 'A-F' | cut -c1-5 \
| xargs -I{} curl -s "https://api.pwnedpasswords.com/range/{}" | head -3
+
+# pad the response with dummy rows so its size tells an observer nothing about the
+# prefix; padded rows always carry a count of 0, so discard them before reading
+curl -s -H 'Add-Padding: true' 'https://api.pwnedpasswords.com/range/5BAA6' \
+ | awk -F: '$2 != 0' | head -3
+
+# NTLM suffixes instead of SHA-1, for checking a dumped AD hash against the corpus.
+# NTLM suffixes are 27 characters where SHA-1 suffixes are 35 — a length mismatch
+# when you diff against a local list means you queried the wrong mode
+curl -s 'https://api.pwnedpasswords.com/range/8846F?mode=ntlm' | head -3
```
-Omitting `user-agent` returns 403 rather than a useful error. A 429 carries a `retry-after`
-header in seconds; honour it, because the limit is per key and hammering it gets the key
-suspended. The output tells you a corpus containing that address was published — not that your
-subject created the account, not that they still use the address, and not what the password was.
+The k-anonymity account endpoint is the one worth building a habit around: it sends six hex
+characters of a SHA-1 and matches the remaining suffix locally, so the address you are
+investigating is never transmitted. The response carries `hashSuffix` and `websites` — not
+`breaches`, which is the field name people assume and then silently get `null` from.
+
+Omitting `user-agent` returns 403 rather than a useful error. `breachedAccount` truncates by
+default, returning only breach names — if your output has no dates in it, you forgot
+`truncateResponse=false`, and a report built on that output cannot say when anything happened. A
+429 carries a `retry-after` header in seconds; honour it, because the limit is per key and
+hammering it gets the key suspended. And the output tells you a corpus containing that address
+was published — not that your subject created the account, not that they still use the address,
+and not what the password was.
### Intelligence X
@@ -462,6 +582,11 @@ What you can assert: a named person, tied by two independent dated documents to
one city, with one address corroborated by a court filing. What you cannot: anything the
aggregator alone said, and any relationship nobody filed.
+What would falsify it: the docket and Form D resolving to different people once a middle name,
+date of birth or signature is compared; the apparent address match being a forwarding or business
+address shared by unrelated parties; or the filing being amended or withdrawn. The company-and-
+city link survives only while the two primary documents point to the same natural person.
+
## Broader catalogues
- [People OSINT](https://tools.osintnewsletter.com/tool-categories/people-osint)
diff --git a/src/content/sheets/osint/social-media-platforms.md b/src/content/sheets/osint/social-media-platforms.md
@@ -4,9 +4,9 @@ description: "What each major platform still exposes to an unauthenticated resea
category: osint
subcategory: "Social Media"
tags: [osint, social-media, telegram, tiktok]
-tools: [yt-dlp, gallery-dl, instaloader, telethon, arctic-shift, atproto]
+tools: [yt-dlp, gallery-dl, instaloader, telethon, telepathy, arctic-shift, atproto, goat, tiktok-hashtag-analysis, instagram-location-search]
difficulty: intermediate
-updated: 2026-09-28
+updated: 2026-10-04
references:
- name: "Bellingcat's Online Investigation Toolkit"
url: "https://bellingcat.gitbook.io/toolkit"
@@ -20,6 +20,36 @@ references:
license: none
relation: derived
note: "Second tool catalogue, cross-checked against the above."
+ - name: "yt-dlp option reference"
+ url: "https://github.com/yt-dlp/yt-dlp#usage-and-options"
+ author: "yt-dlp contributors"
+ relation: link-only
+ note: "Every yt-dlp flag on this page was checked against the project's own option list."
+ - name: "gallery-dl command-line options"
+ url: "https://github.com/mikf/gallery-dl/blob/master/docs/options.md"
+ author: "Mike Fährmann"
+ relation: link-only
+ note: "Every gallery-dl flag on this page was checked against the project's own option list."
+ - name: "Instaloader command-line options"
+ url: "https://instaloader.github.io/cli-options.html"
+ author: "Alexander Graf, André Koch-Kramer"
+ relation: link-only
+ note: "Instaloader flags and target syntax checked against the project's own documentation."
+ - name: "Arctic Shift API reference"
+ url: "https://github.com/ArthurHeitmann/arctic_shift/blob/master/api/README.md"
+ author: "Arthur Heitmann"
+ relation: link-only
+ note: "Endpoint paths and query parameters checked against the project's own API reference."
+ - name: "AT Protocol specifications and lexicons"
+ url: "https://atproto.com/specs/tid"
+ author: "Bluesky Social"
+ relation: link-only
+ note: "XRPC parameter names and the TID timestamp layout checked against the published lexicons."
+ - name: "Meta Ad Library API reference"
+ url: "https://developers.facebook.com/docs/graph-api/reference/ads_archive/"
+ author: "Meta Platforms"
+ relation: link-only
+ note: "ads_archive parameters, enum values and field names checked against the API reference."
---
## What this covers
@@ -49,6 +79,72 @@ change access rules faster than tooling keeps up.
notice. Check the project's last commit before you trust its output, and cross-check one result
by hand.
+The judgement calls: decide early whether you need the content or the network, because bulk media
+capture and graph enumeration use different tools and different amounts of patience. Decide which
+account you are prepared to attribute the work to, before you authenticate anything. And decide
+whether a date is given or derived — if you read a date off the page and then "confirm" it with a
+tool that reads the same page, you have proved nothing.
+
+## The identifier arithmetic
+
+Most platform IDs are not counters. They are clock readings, so the ID you already have dates the
+object with no request to the platform at all. That matters twice over: the object may be deleted
+by the time you look, and the date rendered on a page is frequently the repost's date rather than
+the original's.
+
+**TikTok.** A video ID is a 64-bit value whose top 32 bits hold a Unix timestamp in seconds. Shift
+right by 32 and read it:
+
+```text
+video id : 7169068344567909638
+width : 63 bits, so the timestamp occupies bits 62..32
+id >> 32 : 1669178797
+as UTC : 2022-11-23 04:46:37Z
+```
+
+**X / Twitter.** Status IDs are Snowflakes: 41 bits of milliseconds since a platform-specific
+epoch, sitting 22 bits up. The epoch is the number you have to get right.
+
+```text
+status id : 1949394974262350098
+id >> 22 : 464771979871 milliseconds since the Twitter epoch
++ 1288834974657 : 1753606954528 milliseconds since the Unix epoch
+as UTC : 2025-07-27 09:02:34.528Z
+```
+
+**Bluesky.** A record key is a TID: 13 characters of the sortable base32 alphabet
+`234567abcdefghijklmnopqrstuvwxyz`, decoding to a top bit of 0, then 53 bits of **microseconds**
+since the Unix epoch, then a 10-bit random clock identifier. Because the alphabet sorts, record
+keys sort chronologically as plain strings.
+
+```bash
+# all three, from the shell
+python3 -c 'import datetime as d; i=7169068344567909638; print(d.datetime.fromtimestamp(i>>32, d.timezone.utc))'
+python3 -c 'import datetime as d; i=1949394974262350098; print(d.datetime.fromtimestamp(((i>>22)+1288834974657)/1000, d.timezone.utc))'
+goat syntax tid inspect 3kzifvcppte22
+```
+
+```text
+2022-11-23 04:46:37+00:00
+2025-07-27 09:02:34.528000+00:00
+Timestamp (UTC): 2024-08-12T02:08:03.29Z
+Timestamp (Local): 2024-08-11T19:08:03-07:00
+ClockID: 0
+uint64: 0x187dcbda2b5ca800
+```
+
+Three ways this misleads. The first is the obvious one: an ID dates the **record**, not the
+footage. A clip shot in 2019 and uploaded in 2026 has a 2026 ID, and the arithmetic is still
+correct. The second is who generated the number. TikTok and X mint IDs server-side at creation, so
+they are harder to forge than a rendered date; an atproto TID is minted by whichever component
+writes the record, off that generator's own local clock, and the specification states plainly that
+uniqueness cannot be guaranteed because an adversarial party can create records reusing known TIDs.
+A Bluesky record key is therefore the same class of evidence as `record.createdAt` — an assertion
+by the writer — and the relay's `indexedAt` is the independent number. The third is bit-width drift: quote the shift, not a
+string slice. Taking "the first 31 characters of the binary" works for every current 19-digit
+TikTok ID only because such IDs happen to be 63 bits wide; `id >> 32` keeps working when they are
+not.
+
## Telegram
The most productive platform for open-source research, because public channels are genuinely
@@ -61,8 +157,47 @@ public, history is retained, and forwarding metadata exposes the network between
- **Channel creation dates and message IDs** are sequential, which lets you estimate when a channel
started and how much has been deleted.
-Archiving a channel, mapping its forwards and checking a number against the platform all have
-tooling below under [Key tools](#key-tools).
+Before any tooling, the web preview gives you a channel's recent history with no account, no app
+and no API credentials. It is the cheapest first look there is, and it is the only one that costs
+you nothing in attribution:
+
+```text
+https://t.me/s/channelname the public web view: recent messages, dates, view counts
+https://t.me/channelname/4312 one message, by its sequential id -- the permalink form
+https://t.me/s/channelname?before=4312 page backwards from that id, ~20 messages at a time
+https://t.me/+AbCdEf0123456789 a private invite link; opening it does not join, but
+ it does reveal the chat title and member count
+```
+
+```bash
+# how far back the web view goes, and whether ids are contiguous
+curl -s 'https://t.me/s/channelname' \
+ | grep -o 'data-post="channelname/[0-9]*"' | sed 's/.*\///; s/"//' | sort -n | head -1
+
+# walk backwards in pages of ~20 and collect every id the preview will serve
+for b in 4312 4292 4272 4252; do
+ curl -s "https://t.me/s/channelname?before=$b" \
+ | grep -o 'data-post="channelname/[0-9]*"' | sed 's/.*\///; s/"//'
+done | sort -nu > ids.txt
+
+# the gaps are deletions: compare the id range against the count you actually got
+awk 'NR==1{min=$1} {max=$1; n++} END{print "ids", min"-"max, "present", n, "missing", max-min+1-n}' ids.txt
+```
+
+```text
+ids 401-440 present 39 missing 1
+```
+
+Strip the id with `sed`, not with a second `grep -o '[0-9]*$'`. The attribute ends in a quote, so
+anchoring digits to end-of-line matches the empty string and emits one blank line per post: twenty
+matches in, twenty blanks out, exit status 0. The run looks like a channel with no posts rather
+than like a pipeline that threw the answer away.
+
+The web view is a *preview*, not the archive: it serves a bounded window of recent history and
+will not page back indefinitely, and a channel can disable it entirely. Treat a missing ID as
+evidence of deletion only after you have confirmed the preview is serving that range at all.
+Archiving a channel properly, mapping its forwards and checking a number against the platform all
+have tooling below under [Key tools](#key-tools).
## Facebook
@@ -74,6 +209,30 @@ Graph search is gone and most enumeration is dead. What still works:
open, keyword-searchable, and covers political advertising with spend and reach data.
- **Photo comments and reactions** on public posts still expose participant lists.
+The two URL forms worth memorising. The first resolves a numeric ID that survived a rename; the
+second is the Ad Library filter set, which the web interface writes into its own query string, so
+you can build a link rather than click through four dropdowns:
+
+```text
+# a numeric id back to a profile, which survives every vanity-URL change
+https://www.facebook.com/profile.php?id=100012345678901
+
+# the Ad Library UI's own filter parameters, as it constructs them
+https://www.facebook.com/ads/library/?active_status=all&ad_type=political_and_issue_ads
+ &country=GB&q=climate&search_type=keyword_unordered&media_type=all
+
+# every ad from one page, which is the attribution-relevant view
+https://www.facebook.com/ads/library/?active_status=all&ad_type=all&country=GB
+ &view_all_page_id=123456789
+```
+
+Those are interface parameters observed from the Ad Library itself, not a documented API; the
+documented parameter set is the Graph endpoint under
+[Meta Ad Library API](#meta-ad-library-api) below, and that is what to script against. Everything
+else on Facebook that once had a query form now does not: [Who posted what?](https://whopostedwhat.com/)
+and [Graph Tips](https://graph.tips/) build search URLs for you, and whether they work on any given
+day is a question you answer by trying.
+
## Instagram
- Profile pictures, bios and post counts are visible without login; the feed usually is not.
@@ -82,6 +241,20 @@ Graph search is gone and most enumeration is dead. What still works:
downloads, including comments and geotags where present; it has its own section below.
- Aggressive use gets the account and the IP rate-limited quickly, then blocked outright.
+Instaloader's target syntax is the part worth learning, because it is how you address something
+other than a profile. All four forms are documented targets, not flags:
+
+```bash
+profile_name a public profile, or a private one with --login
+"#hashtag" posts carrying a hashtag; the quotes are usually necessary
+"%212998617" posts tagged with a numeric location id
+"@profile_name" the profiles that profile_name follows, i.e. its followees
+```
+
+Location IDs are the non-obvious pivot: they are numeric, they are stable, and
+[instagram-location-search](#instagram-location-search) below turns a coordinate into a list of
+them, which turns "somewhere near this junction" into a set of addressable targets.
+
## X / Twitter
Heavily restricted since the API changes. Advanced search still works while logged in, and is
@@ -91,6 +264,28 @@ still the best tool on the platform — the operator set is below under
Archived snapshots are now often more reliable than the live site for deleted content — go to the
[Wayback Machine](https://web.archive.org/) first.
+One thing still works with no session at all: a status ID dates itself. If you have a screenshot
+with a URL in it, you have the post's creation time whether or not the post survives:
+
+```bash
+# status id -> UTC creation time, from the Snowflake layout above
+python3 - <<'PY'
+import datetime
+for sid in (1949394974262350098, 1234567890123456789):
+ ms = (sid >> 22) + 1288834974657
+ print(sid, datetime.datetime.fromtimestamp(ms / 1000, datetime.timezone.utc).isoformat())
+PY
+```
+
+```text
+1949394974262350098 2025-07-27T09:02:34.528000+00:00
+1234567890123456789 2020-03-02T19:54:56.824000+00:00
+```
+
+That is a creation time for the ID, which for an original post is the post's time. For a quote-post
+or a reply it is that reply's time, not the thing it replies to, and the `conversation_id` operator
+below is how you get back to the root.
+
## TikTok
- Profiles and individual videos are viewable without an account.
@@ -98,6 +293,31 @@ Archived snapshots are now often more reliable than the live site for deleted co
collects posts by hashtag and charts co-occurrence, which is the useful bulk primitive.
- Video download and metadata extraction both work through `yt-dlp`.
+Dating a TikTok needs neither an account nor a live video, because the ID carries the time. This is
+the whole of what [Bellingcat's extractor](https://bellingcat.github.io/tiktok-timestamp) does, and
+doing it locally means it also works for a URL you only have as text:
+
+```bash
+# pull the id out of the URL and decode it -- no network request at all
+url='https://www.tiktok.com/@user/video/7169068344567909638'
+python3 -c "
+import datetime, re, sys
+vid = int(re.search(r'/video/(\d+)', sys.argv[1]).group(1))
+print(vid, datetime.datetime.fromtimestamp(vid >> 32, datetime.timezone.utc).isoformat())
+" "$url"
+
+# cross-check against what the platform says, which needs the video to still exist
+yt-dlp --dump-json "$url" | jq '{id,timestamp,upload_date,uploader_id,view_count}'
+```
+
+```text
+7169068344567909638 2022-11-23T04:46:37+00:00
+```
+
+When the decoded time and `yt-dlp`'s `timestamp` disagree by more than a second or two, the
+interesting case is a reupload: the ID in the URL belongs to the copy you are looking at, and an
+earlier copy with an earlier ID is what you are actually hunting for.
+
## YouTube
Still the most researcher-friendly large platform: metadata, channel listings, comments and
@@ -105,6 +325,31 @@ auto-generated subtitles all come out of `yt-dlp`, covered below.
Upload timestamps are in the metadata and are a reliable earliest-possible date for footage.
+The two invocations that do most of the work on this platform are a channel listing and a subtitle
+dump — one tells you what exists, the other turns hours of video into text you can grep:
+
+```bash
+# the whole channel as a dated listing, without touching a single video file
+yt-dlp --flat-playlist --dump-json 'https://youtube.com/@channel/videos' \
+ | jq -r '[.upload_date, .id, .duration, .title] | @tsv' | sort
+
+# only uploads inside a window, which is how you tie output to an event
+yt-dlp --flat-playlist --dateafter 20260101 --datebefore 20260331 \
+ -O '%(upload_date)s %(id)s %(title)s' 'https://youtube.com/@channel/videos'
+
+# subtitles for every video in a playlist, as text, no media
+yt-dlp --write-auto-subs --sub-langs 'en.*' --convert-subs srt --skip-download \
+ -o '%(upload_date)s-%(id)s.%(ext)s' 'https://youtube.com/playlist?list=PLID'
+
+# then the actual search, across the lot
+grep -rin 'phrase you care about' *.srt
+```
+
+`--flat-playlist` is what keeps a channel listing cheap: it reads the playlist entries without
+resolving each video, so a 2,000-video channel is one request chain rather than 2,000. The price is
+that some fields are absent from flat entries, so confirm a field exists in `--dump-json` output
+before you build a pipeline on it.
+
## Reddit
- Profiles, posts and comment trees are readable without an account, and appending `.json` to any
@@ -114,6 +359,26 @@ Upload timestamps are in the metadata and are a reliable earliest-possible date
- Subreddit moderator lists, wiki pages and automoderator configs are public and frequently
forgotten by the people who wrote them.
+The `.json` trick, with the parameters that matter:
+
+```bash
+# any Reddit URL plus .json, with escaping turned off and a full page
+curl -s -A 'research-tooling/1.0 (contact: you@example.test)' \
+ 'https://www.reddit.com/user/some_user/comments/.json?limit=100&raw_json=1' \
+ | jq -r '.data.children[] | "\(.data.created_utc) \(.data.subreddit) \(.data.body[0:100])"'
+
+# page with the fullname the previous response handed back
+curl -s -A 'research-tooling/1.0' \
+ 'https://www.reddit.com/r/subreddit/new/.json?limit=100&raw_json=1&after=t3_abc123'
+```
+
+`raw_json=1` is not cosmetic: without it Reddit replaces `<`, `>` and `&` with HTML entities in
+every string field, so a URL or a quoted snippet comes back mangled and your `grep` misses it.
+Expect this route to fail anyway. Reddit now returns **403 to unauthenticated JSON requests from
+datacentre address ranges regardless of User-Agent** — tested from a hosted address, a custom agent
+string makes no difference — so a VPS or CI runner gets nothing while a residential browser
+session gets everything. That asymmetry is exactly why the archive below is not optional.
+
## Bluesky
Structurally the most open platform here, because the AT Protocol is a public API by design: every
@@ -127,6 +392,24 @@ post, follow and profile is a record in a repository you can read without an acc
- **Search is the one thing that is not open.** Post search now needs an authenticated session;
profile and feed reads do not.
+There are two layers and it is worth knowing which one you are on. The AppView
+(`public.api.bsky.app`) serves the rendered, moderated view that the app shows. The account's own
+PDS serves the raw repository, which includes record types the app never displays:
+
+```bash
+# which record collections this account actually holds -- not just posts
+curl -s 'https://bsky.social/xrpc/com.atproto.repo.describeRepo?repo=atproto.com' \
+ | jq '{handle, did, handleIsCorrect, collections}'
+
+# raw records from one collection, straight out of the repository
+curl -s 'https://bsky.social/xrpc/com.atproto.repo.listRecords?repo=atproto.com&collection=app.bsky.feed.post&limit=100&reverse=true' \
+ | jq -r '.records[] | "\(.value.createdAt) \(.uri | split("/") | last) \(.value.text[0:80])"'
+```
+
+`reverse=true` on `listRecords` gives you the **oldest** records first, which is how you read an
+account's first ever posts without paging through everything since. The full tool coverage, the
+AppView endpoints and the `goat` CLI are below under [Key tools](#key-tools).
+
## Key tools
### yt-dlp
@@ -149,9 +432,16 @@ yt-dlp -O '%(upload_date)s %(uploader_id)s %(title)s' 'https://youtube.com/watch
yt-dlp --flat-playlist --dump-json 'https://youtube.com/@channel/videos' \
| jq -r '[.upload_date,.id,.title] | @tsv'
+# the entire playlist as ONE json object, for feeding a script rather than a reader
+yt-dlp --flat-playlist --dump-single-json 'https://youtube.com/@channel/videos' \
+ | jq '{channel: .uploader, count: (.entries | length)}'
+
# auto-generated subtitles, which turn hours of video into greppable text
yt-dlp --write-auto-subs --sub-langs en --skip-download 'https://youtube.com/watch?v=VIDEOID'
+# which subtitle tracks exist at all, before you guess a language code
+yt-dlp --list-subs 'https://youtube.com/watch?v=VIDEOID'
+
# comments, saved alongside the metadata, for identifying participants
yt-dlp --write-info-json --write-comments --skip-download 'https://youtube.com/watch?v=VIDEOID'
@@ -161,21 +451,39 @@ yt-dlp --write-thumbnail --skip-download 'https://youtube.com/watch?v=VIDEOID'
# one clip out of a long stream, by timestamp, instead of the whole six hours
yt-dlp --download-sections '*01:12:30-01:14:00' 'https://youtube.com/watch?v=VIDEOID'
+# a date window across a channel, for tying output to an event
+yt-dlp --flat-playlist --dateafter 20260101 --datebefore 20260331 \
+ -O '%(upload_date)s %(id)s' 'https://youtube.com/@channel/videos'
+
+# filter on any metadata field rather than on dates; "&" is how conditions AND together
+yt-dlp --match-filters 'duration > 600 & view_count > 10000' \
+ -O '%(id)s %(duration)s %(view_count)s' 'https://youtube.com/@channel/videos'
+
+# "?" after the operator keeps items whose field is absent, instead of dropping them silently
+yt-dlp --match-filters 'view_count >? 10000' \
+ -O '%(id)s %(view_count)s' 'https://youtube.com/@channel/videos'
+
# a TikTok video with its metadata — same tool, same flags
yt-dlp --write-info-json 'https://www.tiktok.com/@user/video/1234567890123456789'
# items 1 to 50 of a playlist, with an archive file so a resumed run does not re-fetch
yt-dlp -I 1:50 --download-archive seen.txt 'https://youtube.com/playlist?list=PLID'
+# a livestream from its start, for an event already in progress
+yt-dlp --live-from-start 'https://youtube.com/watch?v=LIVEID'
+
# the logged-in surface, when a public one does not exist — uses your browser's session
yt-dlp --cookies-from-browser firefox --dump-json URL
```
`upload_date` is the platform's own value and is a reliable earliest-possible date for footage;
it is not the date the footage was shot. Fields vary by extractor, so check with `--dump-json`
-before scripting against a key. Passing `--cookies-from-browser` attaches your real session to
-every request, which attributes the activity to that account — use a research profile. Extractors
-break when platforms change, so `pipx upgrade yt-dlp` before blaming the URL.
+before scripting against a key — `--match-filters` silently matches nothing when the field you
+named does not exist on that extractor, which looks identical to a channel with no matching
+videos. The `>?` form above is the documented escape: it treats an absent field as a pass, so a
+sweep that returns nothing is telling you about the videos rather than about the extractor. Passing `--cookies-from-browser` attaches your real session to every request, which
+attributes the activity to that account — use a research profile. Extractors break when platforms
+change, so `pipx upgrade yt-dlp` before blaming the URL.
### gallery-dl
@@ -199,6 +507,9 @@ gallery-dl --write-metadata -D ./capture/username 'https://twitter.com/username/
# direct URLs only, to hand to an archiving tool instead of downloading yourself
gallery-dl -g 'https://www.reddit.com/r/subreddit/comments/abc123/'
+# a chosen field per item, printed as the run goes, instead of parsing json afterwards
+gallery-dl --no-download -N '{date} {num} {filename}' 'https://www.reddit.com/user/username/submitted/'
+
# the first 50 items of a long profile, which is usually all you need to characterise it
gallery-dl --range 1-50 'https://www.flickr.com/photos/username/'
@@ -208,18 +519,29 @@ gallery-dl --filter 'image_width >= 1000' 'https://example-gallery.test/user/x'
# an archive file, so re-running a monitored profile fetches only what is new
gallery-dl --download-archive seen.sqlite3 'https://www.instagram.com/username/'
+# many targets from a file, which is how a watchlist gets run
+gallery-dl -i targets.txt --download-archive seen.sqlite3
+
# slow and polite: the alternative is a 429 and then a block
gallery-dl --sleep 3 --sleep-request 2 --sleep-429 120 'https://www.instagram.com/username/'
+# hash each file as it lands, so the capture is checkable later
+gallery-dl --exec 'sha256sum {} >> hashes.txt' -D ./capture 'https://twitter.com/username/media'
+
+# keep the raw HTML the extractor parsed, which is the only way to debug an empty run
+gallery-dl --write-pages -j 'https://www.instagram.com/username/'
+
# a logged-in session where the public surface is gone
gallery-dl --cookies-from-browser firefox 'https://www.instagram.com/username/'
```
Per-site options live in a config file (`gallery-dl --config-create` writes a starter), and most
-rate-limit problems are solved there rather than on the command line. Timestamps in the sidecar
-JSON are the platform's, which is the point — but the field name differs per extractor, so read
-`-K` output before you build a timeline. A profile capture is a snapshot: items deleted before you
-ran it are simply absent, and nothing in the output tells you that.
+rate-limit problems are solved there rather than on the command line — `-o KEY=VALUE` sets one
+inline when you do not want to edit the file. Timestamps in the sidecar JSON are the platform's,
+which is the point — but the field name differs per extractor, so read `-K` output before you build
+a timeline. A profile capture is a snapshot: items deleted before you ran it are simply absent, and
+nothing in the output tells you that. An extractor that returns zero items is far more often broken
+than correct, and `--write-pages` is how you tell the difference.
### Instaloader
@@ -241,62 +563,152 @@ instaloader --comments --geotags profile_name
# just the profile picture at full resolution, for reverse image search
instaloader --no-posts profile_name
+# cap a location run; --count does not apply to ordinary profile targets
+instaloader --count 50 --no-videos '%212998617'
+
# posts within a date window, using a Python filter expression
instaloader --post-filter 'date_utc >= datetime(2026,1,1)' profile_name
+# only posts that carry a location, which is the geolocation-relevant subset
+instaloader --post-filter 'location is not None' --geotags profile_name
+
# a hashtag rather than an account
instaloader '#hashtag'
+# a numeric location id, which is what instagram-location-search below produces
+instaloader '%212998617'
+
+# everywhere this account has been tagged by other people
+instaloader --tagged --no-posts profile_name
+
+# one post by shortcode; the leading dash needs the -- separator first
+instaloader -- -B_Nmp6MlQGV
+
+# readable JSON instead of xz-compressed, for grepping a capture afterwards
+instaloader --no-compress-json --no-pictures --no-videos profile_name
+
# stories and highlights need a session; both attribute the view to that account
instaloader --login YOUR_RESEARCH_ACCOUNT --stories --highlights profile_name
-# resume a monitored profile from where you stopped
-instaloader --fast-update --latest-stamps stamps.ini profile_name
+# reuse a browser session instead of handing credentials to the tool
+instaloader --load-cookies firefox --sessionfile ./research.session profile_name
+
+# stop on the first rate-limit response rather than grinding into a block
+instaloader --abort-on 429,401 --max-connection-attempts 1 profile_name
+
+# resume by stopping at the first already-downloaded post
+instaloader --fast-update profile_name
+
+# resume by recorded timestamp instead, which is the only resume a metadata-only run can use
+instaloader --no-pictures --no-videos --latest-stamps stamps.ini profile_name
```
**Viewing a story is attributed.** The account whose session you used appears in the poster's
viewer list, so `--stories` is never a passive operation. Unauthenticated use is rate-limited
aggressively and then blocked by IP; authenticated use gets the account flagged and sometimes
-disabled. Geotags are only present where the poster added them, and the location name is
-user-chosen, not GPS.
+disabled — `--abort-on 429` is the difference between one bad response and a dead account.
+`--no-pictures` cannot be combined with `--fast-update`, because the resume logic keys off
+downloaded media, so a metadata-only monitoring run needs `--latest-stamps` instead. Geotags are
+only present where the poster added them, and the location name is user-chosen, not GPS.
+
+### instagram-location-search
+
+Bellingcat's answer to "what is Instagram's own name for this place". It takes a coordinate and
+returns the location tags Instagram knows about nearby, with their numeric IDs — which is the
+pivot, because an ID is an addressable target for Instaloader while a place name is not.
+
+```bash
+pip install instagram-location-search
+
+# a coordinate to a list of location tags, as CSV
+instagram-location-search --cookie "$IG_COOKIE" --lat 51.5074 --lng -0.1278 --csv locs.csv
+
+# every output format at once: the raw API response, a geospatial layer, and a map to eyeball
+instagram-location-search --cookie "$IG_COOKIE" --lat 51.5074 --lng -0.1278 \
+ --json locs.json --geojson locs.geojson --map locs.html
+
+# just the IDs, which is the form the next tool wants
+instagram-location-search --cookie "$IG_COOKIE" --lat 51.5074 --lng -0.1278 --ids locs.txt
+
+# then feed each one to Instaloader as a location target
+while read -r id; do instaloader --no-videos --geotags "%$id"; done < locs.txt
+```
+
+It needs a session cookie, so every search is attributed to that account. The locations it returns
+are *Instagram's* places, which are user-created: there are duplicates, misspellings, joke entries
+and venues that have closed, and a radius search returns them by proximity to the coordinate rather
+than by relevance. A location tag on a post means the poster selected that place from a list, not
+that a GPS fix put them there — see [Geolocation & Chronolocation](/sheets/osint/geolocation) for
+what actually constitutes placing an image.
### Telegram: Telepathy and Telethon
[Telepathy](https://github.com/proseltd/Telepathy-Community) is the toolkit the Telegram sections
-of every OSINT guide point at, and it still runs — but it has been **unmaintained since July
-2024**, with 2.3.4 the last release and the maintainers saying so in the README. Use it knowing
-that; for anything you need to keep working, the durable route is Telethon, which tracks the
-Telegram API itself.
+of older OSINT guides point at, but it has been **unmaintained since July 2024** and the current
+source branch does not install cleanly. The commands below document the published 2.3.4 interface;
+for anything you need to keep working, the durable route is Telethon, which tracks the Telegram
+API itself.
```bash
-pipx install telepathy
+pip3 install telepathy
-telepathy -t channelname # basic channel scrape
-telepathy -t channelname -c # comprehensive: messages, members, media
+telepathy -t channelname # basic channel scan
+telepathy -t channelname -c # comprehensive: adds the message history archive
telepathy -t channelname -c -f # plus a forward edgelist, the useful part
-telepathy -u username # look up a user
-telepathy -e # export your account's chat list
+telepathy -t channelname -c -m # media archiving alongside the comprehensive scan
+telepathy -t channelname -c -r # archive channel replies as well as posts
+telepathy -t username -u # look up a user your account has encountered
+telepathy -e # export your account's chat list to CSV
+telepathy -t channelname -a 2 # run from an alternative number or API key set
```
Both need API credentials from [my.telegram.org](https://my.telegram.org/), which are tied to a
-real phone number — use one you are willing to lose. Telethon in twenty lines does the part that
+real phone number — use one you are willing to lose. Telethon in forty lines does the part that
matters, and does not rot:
```python
# pip install telethon
+from collections import Counter
from telethon.sync import TelegramClient
+from telethon.errors import FloodWaitError
+from telethon.tl.functions.channels import GetFullChannelRequest
+from telethon.tl.types import InputMessagesFilterPhotos
with TelegramClient('research-session', API_ID, API_HASH) as client:
entity = client.get_entity('channelname')
print(entity.id, entity.title, entity.date) # channel id and creation date
+ # what Telegram itself reports about the channel, including the linked discussion group
+ full = client(GetFullChannelRequest(entity)).full_chat
+ print(full.participants_count, full.linked_chat_id, (full.about or '')[:120])
+
# message history, newest first; message IDs are sequential, so gaps are deletions
for msg in client.iter_messages(entity, limit=500):
fwd = msg.forward.chat.username if msg.forward and msg.forward.chat else None
print(msg.id, msg.date.isoformat(), fwd or '-', (msg.message or '')[:80])
+ # the count Telegram reports, against the newest id: the difference is deleted content
+ page = client.get_messages(entity, limit=100)
+ newest = page[0]
+ print('reported total', page.total, 'newest id', newest.id, 'gap', newest.id - page.total)
+
+ # oldest first, which is how you read a channel's first posts without paging the lot
+ for msg in client.iter_messages(entity, limit=20, reverse=True):
+ print(msg.id, msg.date.isoformat(), (msg.message or '')[:80])
+
+ # server-side keyword search inside one channel -- far cheaper than fetching everything
+ for msg in client.iter_messages(entity, search='keyword', limit=100):
+ print(msg.id, msg.date.isoformat(), (msg.message or '')[:120])
+
+ # only messages carrying photos, for building a verification set
+ for msg in client.iter_messages(entity, filter=InputMessagesFilterPhotos, limit=50):
+ print(msg.id, msg.date.isoformat(), msg.file.name if msg.file else '-')
+
+ # everything one account posted in a group, which is the behavioural view
+ for msg in client.iter_messages('groupname', from_user='username', limit=200):
+ print(msg.id, msg.date.isoformat(), (msg.message or '')[:80])
+
# the forward graph: which channels this one amplifies, with counts
- from collections import Counter
sources = Counter(
m.forward.chat.username
for m in client.iter_messages(entity, limit=2000)
@@ -310,9 +722,11 @@ with TelegramClient('research-session', API_ID, API_HASH) as client:
```
Sequential message IDs are the quiet win here: a gap between IDs is deleted content, and the first
-ID dates the channel. Joining a group to read it adds your research account to a list other members
-can see, and scraping at speed from one account is exactly the behaviour Telegram bans for. Rate
-limits surface as `FloodWaitError` with a duration — sleep for it rather than retrying.
+ID dates the channel. The `search=` parameter runs server-side, which matters for a channel with
+100,000 messages — it is the difference between one request and a thousand. Joining a group to read
+it adds your research account to a list other members can see, and scraping at speed from one
+account is exactly the behaviour Telegram bans for. Rate limits surface as `FloodWaitError` with a
+`seconds` attribute — sleep for exactly that, rather than retrying, because retrying extends it.
### telegram-phone-number-checker
@@ -337,6 +751,10 @@ telegram-phone-number-checker --usernames johndoe
# both in one run, to a named output file
telegram-phone-number-checker --phone-numbers +14155550123 --usernames johndoe --output results.json
+# credentials on the command line, overriding the .env, for a one-off run on a second account
+telegram-phone-number-checker --api-id 123456 --api-hash abcdef0123456789 \
+ --api-phone-number +14155550000 --phone-numbers +14155550123
+
# through a proxy, since the lookups come from your account
pip install 'telegram-phone-number-checker[proxy]'
telegram-phone-number-checker --phone-numbers +14155550123 --proxy socks5://127.0.0.1:1080
@@ -372,6 +790,12 @@ tiktok-hashtag-analysis london --number 20 --plot
# download the videos too, capped per hashtag
tiktok-hashtag-analysis london --download --limit 200
+# short forms, for the same thing: -d download, -t table, -p plot
+tiktok-hashtag-analysis london -d -t --number 20 --limit 50
+
+# keep a log file, which is the only record of what a long run actually did
+tiktok-hashtag-analysis london --log ./run.log --output-dir ./tiktok-run
+
# watch the browser work, which is how you debug a run that returns nothing
tiktok-hashtag-analysis london --headed -v
```
@@ -379,9 +803,9 @@ tiktok-hashtag-analysis london --headed -v
It drives a real browser, so it is slow and it breaks when TikTok changes its front end — a run
that returns zero posts for a busy hashtag means the scraper is broken, not that the hashtag is
empty. Check with `--headed` before concluding anything. Hashtag collection is a sample, never the
-complete set, so counts are comparative at best. For a single video's exact upload time,
-[Bellingcat's TikTok timestamp extractor](https://bellingcat.github.io/tiktok-timestamp) decodes it
-from the video ID, and `yt-dlp --dump-json` gets the rest.
+complete set, so counts are comparative at best, and two runs an hour apart will not agree. For a
+single video's exact upload time, decode the ID as shown under [TikTok](#tiktok) above, and
+`yt-dlp --dump-json` gets the rest.
### Arctic Shift
@@ -405,6 +829,10 @@ curl -s -G "$AS/comments/search" --data-urlencode 'author=some_user' \
curl -s -G "$AS/posts/search" --data-urlencode 'subreddit=worldnews' \
--data-urlencode 'title=wuhan' --data-urlencode 'after=2019-12-30' --data-urlencode 'limit=10'
+# title and selftext together, which is what 'query' is for
+curl -s -G "$AS/posts/search" --data-urlencode 'subreddit=osint' \
+ --data-urlencode 'query=telegram' --data-urlencode 'limit=25'
+
# only the fields you want, which keeps a wide sweep manageable
curl -s -G "$AS/posts/search" --data-urlencode 'subreddit=osint' \
--data-urlencode 'fields=title,created_utc,author' --data-urlencode 'limit=100'
@@ -413,19 +841,55 @@ curl -s -G "$AS/posts/search" --data-urlencode 'subreddit=osint' \
curl -s -G "$AS/posts/search" --data-urlencode 'url=example-news-daily.com' \
--data-urlencode 'limit=100' | jq -r '.data[].subreddit' | sort | uniq -c | sort -rn
+# posting cadence rather than content: a month-by-month count
+curl -s -G "$AS/posts/search/aggregate" --data-urlencode 'subreddit=osint' \
+ --data-urlencode 'aggregate=created_utc' --data-urlencode 'frequency=month' \
+ --data-urlencode 'after=2026-01-01' | jq -r '.data[] | "\(.created_utc[0:7]) \(.count)"'
+
+# where an account actually operates, weighted and ranked
+curl -s -G "$AS/users/interactions/subreddits" --data-urlencode 'author=some_user' \
+ --data-urlencode 'limit=10' | jq -r '.data[] | "\(.count)\t\(.subreddit)"'
+
+# who an account talks to, which is the social graph Reddit does not expose
+curl -s -G "$AS/users/interactions/users" --data-urlencode 'author=some_user' \
+ --data-urlencode 'min_count=3' --data-urlencode 'limit=20'
+
+# every wiki page a subreddit has, including the ones nobody links to
+curl -s -G "$AS/subreddits/wikis/list" --data-urlencode 'subreddit=osint' | jq -r '.data[]'
+
# the full comment tree under one post, up to 25,000 comments
-curl -s -G "$AS/comments/tree" --data-urlencode 'link_id=abc123'
+curl -s -G "$AS/comments/tree" --data-urlencode 'link_id=abc123' \
+ --data-urlencode 'limit=9999'
# known IDs straight to records, up to 500 per call
curl -s -G "$AS/comments/ids" --data-urlencode 'ids=t1_aaa,t1_bbb'
+
+# accounts by username prefix, for a sockpuppet family
+curl -s -G "$AS/users/search" --data-urlencode 'author_prefix=civicwatch' \
+ --data-urlencode 'sort_type=author' --data-urlencode 'limit=100' | jq -r '.data[].author'
+
+# an RSS feed of a standing query, for monitoring rather than a one-off sweep
+curl -s -G "$AS/comments/search" --data-urlencode 'subreddit=osint' \
+ --data-urlencode 'format=rss' --data-urlencode 'limit=25'
+```
+
+A real response, from the aggregate query above:
+
+```json
+{"data":[{"created_utc":"2025-12-31T23:00:00.000Z","count":"390"},
+ {"created_utc":"2026-01-31T23:00:00.000Z","count":"376"},
+ {"created_utc":"2026-02-28T23:00:00.000Z","count":"503"},
+ {"created_utc":"2026-03-31T22:00:00.000Z","count":"450"}]}
```
-Keyword search on `title`, `selftext` or `body` requires an accompanying `author`, `subreddit`,
-`link_id` or `parent_id` — a bare keyword sweep is refused, and it is refused for very active
-authors and subreddits too. The archive holds what it ingested at the time, so an edit after
-ingestion is invisible and a post removed within seconds of being made may never have been
-captured. A record here is evidence that text was published, not that it is still live, and the
-live check is a separate step.
+Note the bucket boundaries: `23:00:00Z` then `22:00:00Z` after March, because the buckets are cut
+in a local timezone that observes DST, so a month's `created_utc` label is the start of the bucket
+and not the month. Keyword search is constrained by design — `title` works only alongside `author`
+or `subreddit`, `body` only alongside `author`, `subreddit`, `link_id` or `parent_id` — so a bare
+keyword sweep is refused, and it is refused for very active authors and subreddits too. The archive
+holds what it ingested at the time, so an edit after ingestion is invisible and a post removed
+within seconds of being made may never have been captured. A record here is evidence that text was
+published, not that it is still live, and the live check is a separate step.
### Bluesky and the AT Protocol
@@ -447,6 +911,11 @@ curl -s 'https://plc.directory/did:plc:z72i7hdynmk6r22z27h6tvur/log/audit' \
curl -s "$API/app.bsky.actor.getProfile?actor=bellingcat.com" \
| jq '{handle,displayName,followersCount,postsCount,createdAt}'
+# up to 25 accounts in one request, which is how you compare a suspected cluster
+curl -s -G "$API/app.bsky.actor.getProfiles" \
+ --data-urlencode 'actors=bellingcat.com' --data-urlencode 'actors=atproto.com' \
+ | jq -r '.profiles[] | "\(.createdAt) \(.followersCount)\t\(.handle)"'
+
# the author's posts, paged with the cursor the response returns
curl -s "$API/app.bsky.feed.getAuthorFeed?actor=bellingcat.com&limit=100" \
| jq -r '.feed[] | "\(.post.indexedAt) \(.post.record.text[0:100])"'
@@ -454,23 +923,113 @@ curl -s "$API/app.bsky.feed.getAuthorFeed?actor=bellingcat.com&limit=100" \
# only posts carrying media, which is the usual filter for verification work
curl -s "$API/app.bsky.feed.getAuthorFeed?actor=bellingcat.com&limit=50&filter=posts_with_media"
+# the account's own threads without the replies noise
+curl -s "$API/app.bsky.feed.getAuthorFeed?actor=bellingcat.com&limit=50&filter=posts_no_replies"
+
# the follow graph, fully enumerable in both directions
curl -s "$API/app.bsky.graph.getFollows?actor=bellingcat.com&limit=100" | jq -r '.follows[].handle'
curl -s "$API/app.bsky.graph.getFollowers?actor=bellingcat.com&limit=100" | jq -r '.followers[].handle'
+# the lists an account curates, which say who it considers relevant
+curl -s "$API/app.bsky.graph.getLists?actor=bellingcat.com&limit=50" \
+ | jq -r '.lists[] | "\(.purpose)\t\(.listItemCount)\t\(.name)"'
+
+# who amplified one post, and who liked it -- both need the post's AT-URI
+curl -s "$API/app.bsky.feed.getRepostedBy?uri=at://did:plc:xxxx/app.bsky.feed.post/yyyy&limit=100" \
+ | jq -r '.repostedBy[].handle'
+curl -s "$API/app.bsky.feed.getLikes?uri=at://did:plc:xxxx/app.bsky.feed.post/yyyy&limit=100" \
+ | jq -r '.likes[] | "\(.createdAt) \(.actor.handle)"'
+
# a whole thread, replies included, from one post's AT URI
curl -s "$API/app.bsky.feed.getPostThread?uri=at://did:plc:xxxx/app.bsky.feed.post/yyyy"
-# post search needs a session — this returns 403 on the public host
-curl -s "$API/app.bsky.feed.searchPosts?q=osint&limit=5"
+# post search needs a session — this returns an auth error on the public host
+curl -s -G "$API/app.bsky.feed.searchPosts" --data-urlencode 'q=osint' \
+ --data-urlencode 'sort=latest' --data-urlencode 'since=2026-01-01' --data-urlencode 'limit=25'
+```
+
+Once you have a session, `searchPosts` is the richest query surface on the platform: `q` plus
+`author`, `mentions`, `domain`, `url`, `lang`, repeated `tag` parameters that AND together, and
+`since`/`until` which take a bare `YYYY-MM-DD`. Authenticate with
+`com.atproto.server.createSession` against the account's own PDS and send the returned
+`accessJwt` as a bearer token:
+
+```bash
+# a session on the account's own PDS; use an app password, never the real one
+PDS='https://bsky.social'
+TOKEN=$(curl -s -X POST "$PDS/xrpc/com.atproto.server.createSession" \
+ -H 'Content-Type: application/json' \
+ -d '{"identifier":"you.bsky.social","password":"xxxx-xxxx-xxxx-xxxx"}' | jq -r .accessJwt)
+
+# every post from one account mentioning a domain, inside a date window
+curl -s -G "$PDS/xrpc/app.bsky.feed.searchPosts" -H "Authorization: Bearer $TOKEN" \
+ --data-urlencode 'q=*' --data-urlencode 'author=bellingcat.com' \
+ --data-urlencode 'domain=example-news-daily.com' --data-urlencode 'since=2026-01-01' \
+ | jq -r '.posts[] | "\(.record.createdAt) \(.author.handle) \(.record.text[0:100])"'
+
+# posts carrying two tags at once, which is the coordination signal
+curl -s -G "$PDS/xrpc/app.bsky.feed.searchPosts" -H "Authorization: Bearer $TOKEN" \
+ --data-urlencode 'q=*' --data-urlencode 'tag=osint' --data-urlencode 'tag=verification' \
+ --data-urlencode 'sort=latest' | jq -r '.posts[].author.handle' | sort | uniq -c | sort -rn
```
The PLC audit log is the find here: it is an append-only record of every handle the account has
used and every server it has moved between, dated, with the first entry establishing when the
account was created. `indexedAt` is when the relay saw the post, not when the author wrote it —
-`record.createdAt` is the author's claim and is trivially spoofable, so quote both. For search,
-authenticate with `com.atproto.server.createSession` against the account's own PDS and send the
-returned access token as a bearer; that turns a keyless read into an attributable one.
+`record.createdAt` is the author's claim and is trivially spoofable, as is the TID in the record
+key, so quote both and say which is which. `searchPosts` is explicitly documented as possibly
+non-public, so treat the public host returning results for it as luck rather than a contract, and
+note that `since`/`until` filter on the server's `sortAt`, which may not match `createdAt` at all.
+Authenticating turns a keyless read into an attributable one.
+
+### goat
+
+The reference AT Protocol CLI, from Bluesky themselves. It does the things `curl` makes awkward:
+exporting a whole repository as a single file, decoding record keys, walking the PLC operation log,
+and tailing the firehose. For Bluesky work it replaces about ten of the commands above.
+
+```bash
+brew install goat # or: go install github.com/bluesky-social/goat@latest
+
+# resolve an identity to its DID document, including which PDS holds the data
+goat resolve wyden.senate.gov
+
+# every record collection the account holds -- including ones the app never renders
+goat ls -c dril.bsky.social
+
+# one record, as the author wrote it
+goat get at://dril.bsky.social/app.bsky.feed.post/3kkreaz3amd27
+
+# the entire repository as a CAR file: every record, in one evidential artefact
+goat repo export jay.bsky.team
+
+# and the blobs, which is every image the account has ever uploaded
+goat blob export jay.bsky.team
+
+# the full identity history: handle changes, PDS moves, key rotations, dated
+goat plc history atproto.com
+
+# a record key back to the time the client claims it was written
+goat syntax tid inspect 3kzifvcppte22
+
+# live posts matching nothing in particular, for watching an event unfold
+goat firehose --ops -c app.bsky.feed.post
+
+# handle changes across the whole network, which is a renaming-detection feed
+goat firehose --account-events | jq .payload.handle
+
+# any XRPC query against any host, which is the escape hatch
+goat xrpc query https://public.api.bsky.app app.bsky.actor.getProfile actor==atproto.com
+```
+
+`goat repo export` is the one to reach for when the account might be deleted: a CAR file is the
+whole repository at one moment, signed, and it keeps working after the account does not. Note the
+`==` in the `xrpc query` form — that is how `goat` distinguishes a query parameter from a header,
+and a single `=` there means something else. Reads need no authentication; `goat account login`
+exists for writes and for the handful of authenticated reads, and logging in attributes everything
+after it. A TID decoded by `syntax tid inspect` reads the clock of whatever generated the record
+key, not the network's, and the TID specification notes that known TIDs can be reused — so it
+dates a claim rather than an event.
### Meta Ad Library API
@@ -488,6 +1047,11 @@ curl -s -G "$G" -d "access_token=$TOKEN" \
-d 'search_terms=climate' -d "ad_reached_countries=['GB']" \
-d 'fields=page_name,ad_delivery_start_time,ad_snapshot_url' | jq -r '.data[].page_name' | sort -u
+# an exact phrase rather than unordered keywords
+curl -s -G "$G" -d "access_token=$TOKEN" \
+ -d 'search_terms=net zero' -d 'search_type=KEYWORD_EXACT_PHRASE' \
+ -d "ad_reached_countries=['GB']" -d 'fields=page_name,ad_creative_bodies'
+
# one page's entire ad history, which is the attribution-relevant query
curl -s -G "$G" -d "access_token=$TOKEN" \
-d 'search_page_ids=123456789' -d "ad_reached_countries=['GB']" \
@@ -497,29 +1061,51 @@ curl -s -G "$G" -d "access_token=$TOKEN" \
# who paid, which is the field the transparency rules exist for
curl -s -G "$G" -d "access_token=$TOKEN" \
-d 'search_terms=election' -d "ad_reached_countries=['GB']" \
- -d 'fields=bylines,page_name,spend,currency,publisher_platforms'
+ -d 'fields=bylines,beneficiary_payers,page_name,spend,currency,publisher_platforms'
# where it ran and to whom
curl -s -G "$G" -d "access_token=$TOKEN" \
-d 'search_terms=election' -d "ad_reached_countries=['GB']" \
-d 'fields=delivery_by_region,demographic_distribution,languages'
+# only the ads that ran as video, on Instagram rather than Facebook
+curl -s -G "$G" -d "access_token=$TOKEN" \
+ -d 'search_terms=election' -d "ad_reached_countries=['GB']" \
+ -d 'media_type=VIDEO' -d "publisher_platforms=['INSTAGRAM']" \
+ -d 'fields=page_name,ad_snapshot_url,publisher_platforms'
+
+# the non-political categories, which most people never query
+curl -s -G "$G" -d "access_token=$TOKEN" \
+ -d 'search_terms=apartment' -d "ad_reached_countries=['GB']" \
+ -d 'ad_type=HOUSING_ADS' -d 'fields=page_name,ad_creative_bodies,ad_delivery_start_time'
+
# a date window, for tying a campaign to an event
curl -s -G "$G" -d "access_token=$TOKEN" \
-d 'search_terms=referendum' -d "ad_reached_countries=['IE']" \
-d 'ad_delivery_date_min=2026-01-01' -d 'ad_delivery_date_max=2026-03-31' \
-d 'fields=page_name,spend,ad_snapshot_url'
+# reach, not impressions: the EU and Brazil figures are per-ad totals
+curl -s -G "$G" -d "access_token=$TOKEN" \
+ -d 'search_page_ids=123456789' -d "ad_reached_countries=['IE']" \
+ -d 'fields=eu_total_reach,age_country_gender_reach_breakdown,estimated_audience_size'
+
# page the result set with the cursor Graph hands back
curl -s -G "$G" -d "access_token=$TOKEN" -d 'search_terms=climate' \
-d "ad_reached_countries=['GB']" -d 'fields=page_name' -d 'limit=100' -d 'after=CURSOR'
```
Leaving out both `search_terms` and `search_page_ids` is a parameter error, not an unfiltered
-dump. `search_terms` matches the ad's own language and is capped at 100 characters, and spaces are
-treated as AND. `spend` and `impressions` come back as ranges, not numbers — report them as
-ranges. Coverage is political and issue ads in the countries where Meta is required to disclose
-them; commercial ads appear in the web Ad Library with far less metadata and are not in the API
+dump. `search_terms` matches the ad's own language and is capped at 100 characters, spaces are
+treated as AND unless you set `search_type=KEYWORD_EXACT_PHRASE`, and `search_page_ids` takes at
+most ten IDs. `spend` and `impressions` come back as ranges, not numbers — report them as ranges,
+and note that `eu_total_reach` is a single figure while `estimated_audience_size` is another range.
+Several fields and filters — `bylines`, `delivery_by_region`, `demographic_distribution`,
+`estimated_audience_size_min`/`_max` — exist only for political and issue ads, so an empty result
+for a commercial advertiser means the field does not apply rather than that the ad does not exist.
+Coverage is political and issue ads in the countries where Meta is required to disclose them, plus
+the `EMPLOYMENT_ADS`, `FINANCIAL_PRODUCTS_AND_SERVICES_ADS` and `HOUSING_ADS` categories;
+ordinary commercial ads appear in the web Ad Library with far less metadata and are not in the API
the same way.
### X / Twitter advanced search
@@ -533,18 +1119,49 @@ from:username since:2026-01-01 until:2026-06-30 one account, bounded by date
to:username replies directed at an account
"exact phrase" -filter:retweets the original post rather than its echoes
filter:images filter:links posts carrying media or outbound links
+filter:replies only replies, which is where arguments live
url:example-news-daily.com who linked a domain
geocode:51.5074,-0.1278,5km posts placed near a coordinate
-min_faves:500 min_retweets:100 the versions that actually travelled
+min_faves:500 min_retweets:100 min_replies:50 the versions that actually travelled
+lang:ru "phrase" one language, for a multilingual campaign
+(term1 OR term2) -term3 grouping and negation, as you would expect
list:12345678 "phrase" search inside a list you have built
conversation_id:1234567890123456789 one thread, including replies
+from:username filter:media since:2026-03-01 one account's images in one month
```
Every one of these requires a logged-in session, which attributes the searching to that account —
use a research account, and assume the queries are logged. Date-bounded `from:` queries are the
-most reliable form; free-text search silently drops older results. For deleted posts, go to the
-[Wayback Machine](https://web.archive.org/) and archive.today first: for X specifically, archived
-snapshots are now a more dependable source than the live site.
+most reliable form; free-text search silently drops older results, so an empty result for an old
+date range is not evidence of absence. X publishes no machine-readable operator list and its help
+pages block automated fetches, so verify any operator against the advanced-search form itself
+before relying on it — the form is the only authority, and it has quietly lost operators before.
+For deleted posts, go to the [Wayback Machine](https://web.archive.org/) and archive.today first:
+for X specifically, archived snapshots are now a more dependable source than the live site. A
+status ID still dates itself with no session at all, as shown under [X / Twitter](#x--twitter)
+above.
+
+## Platform tooling at a glance
+
+| Tool | Platform | Account needed | What it gets you |
+| --- | --- | --- | --- |
+| [yt-dlp](https://github.com/yt-dlp/yt-dlp) | 1,000+ sites | No, until the platform insists | Video and channel metadata, subtitles, comments, timed clips |
+| [gallery-dl](https://github.com/mikf/gallery-dl) | Images and galleries, most platforms | No, until the platform insists | Bulk profile capture with per-item JSON sidecars |
+| [Instaloader](https://instaloader.github.io/) | Instagram | No for public posts, yes for stories | Posts, captions, comments, geotags, profile pictures |
+| [instagram-location-search](https://github.com/bellingcat/instagram-location-search) | Instagram | Yes, a session cookie | Location IDs near a coordinate, as CSV, JSON, GeoJSON or a map |
+| [Telethon](https://docs.telethon.dev/) | Telegram | Yes, API credentials and a phone number | Message history, forward graph, member lists, creation dates |
+| [Telepathy](https://github.com/proseltd/Telepathy-Community) | Telegram | Yes, API credentials and a phone number | The same as a CLI, but unmaintained since July 2024 |
+| [Telegram Phone Number Checker](https://github.com/bellingcat/telegram-phone-number-checker) | Telegram | Yes, and it writes to your contacts | Whether a number or username is registered |
+| [tiktok-hashtag-analysis](https://github.com/bellingcat/tiktok-hashtag-analysis) | TikTok | No | Posts by hashtag, plus a co-occurrence table or plot |
+| [Arctic Shift](https://arctic-shift.photon-reddit.com/) | Reddit | No | Deleted posts and comments, wiki page lists, activity aggregates |
+| [goat](https://github.com/bluesky-social/goat) | Bluesky / AT Protocol | No for reads | Repository and blob exports, PLC history, TID decoding, firehose |
+| [Meta Ad Library API](https://developers.facebook.com/docs/graph-api/reference/ads_archive/) | Facebook, Instagram | Yes, an approved developer app | Spend and reach ranges, payer bylines, delivery breakdowns |
+| [X/Twitter Advanced Search](https://x.com/search-advanced) | X | Yes, a logged-in session | Date-, geo- and engagement-bounded post search |
+
+Reach for `yt-dlp` and `gallery-dl` first on any platform either covers, because they are the only
+two on this list maintained at the pace platforms break things. Reach for Arctic Shift and `goat`
+when the question is historical, because both hold data the live platform will not serve. Everything
+else here is worth a check against its last commit date before you trust its output.
## Tool reference
@@ -610,6 +1227,40 @@ snapshots are now a more dependable source than the live site.
Check a tool's commit history before trusting its output.
- **Deleted is not gone, and present is not original.** Archive first, then analyse.
- **Reposts dominate.** The account with the most engagement is rarely the origin. Trace back.
+- **An ID dates the record, not the footage.** Decoding a TikTok or Snowflake ID tells you when
+ that upload was created. A 2019 clip uploaded in 2026 has a 2026 ID and the arithmetic is still
+ right.
+- **A client-written timestamp is a claim.** An atproto TID and `record.createdAt` are both chosen
+ by whatever wrote the record; `indexedAt` is the relay's. Quote the one the account did not
+ control, and say which you are quoting.
+- **Zero results and broken scraper look identical.** A busy hashtag returning no posts, or a
+ `--match-filters` expression naming a field that extractor does not emit, both print nothing.
+ Run one query you already know the answer to before you believe a negative.
+- **A pipeline can match and then discard the match.** `grep -o '[0-9]*$'` against
+ `data-post="channel/441"` anchors digits to a line that ends in a quote, so it emits a blank
+ line per post and exits 0. The extractor worked, the shell threw the answer away, and the
+ symptom is an empty file. Count the lines at each stage of a new pipeline before trusting the
+ last one.
+- **A date filter may not filter on the date you mean.** Bluesky's `since`/`until` on
+ `searchPosts` are documented against the server's `sortAt`, which need not equal
+ `record.createdAt`; Arctic Shift's `created_utc` aggregate buckets are cut in a DST-observing
+ local zone. Both return a defensible answer to a question you did not ask.
+- **Unauthenticated does not mean unattributed.** Your IP is an identifier, and a sweep from one
+ address is a signature. Conversely, authenticating to get past a block converts an anonymous read
+ into a logged one tied to an account you will lose.
+- **Datacentre addresses are blocked where browsers are not.** Reddit's `.json` surface returns 403
+ from hosted ranges regardless of User-Agent. Code that works on a laptop and fails in CI has not
+ found a bug; it has found a policy.
+- **Archives ingest, they do not watch.** Arctic Shift holds what it captured when it captured it:
+ an edit made afterwards is invisible, and something deleted within seconds may never have been
+ recorded at all.
+- **Spend and reach are ranges.** Meta reports `spend` and `impressions` as buckets. Writing a
+ midpoint and calling it a figure invents precision the source does not have.
+- **Counts from a sample are comparative at best.** Hashtag collection, search results and "top
+ twenty" lists are samples whose size you do not control. Two runs an hour apart will disagree,
+ which is a property of the method, not of the world.
+- **The same handle on four platforms is one lead, not one person.** It takes content, timing or a
+ payment record to tie accounts together. Say which of the three you have.
## Worked example
@@ -620,31 +1271,57 @@ attached and no other context.
[username and account discovery](/sheets/osint/usernames-and-accounts) — returns hits on
Bluesky, YouTube, Reddit and a Telegram channel of the same name. Four surfaces, four different
amounts of openness.
-2. **Date the accounts, keylessly.** `resolveHandle` gives the Bluesky DID; the PLC audit log's
- first entry dates the account to March 2026 and shows one earlier handle. An account presenting
- itself as an established watchdog, three months old, is the first real finding.
+2. **Date the accounts, keylessly.** `resolveHandle` gives the Bluesky DID
+ `did:plc:z72i7hdynmk6r22z27h6tvur`; the PLC audit log's first entry dates the account to
+ **12 March 2026** and shows one earlier handle. An account presenting itself as an established
+ watchdog, three months old, is the first real finding.
3. **Characterise the output.** `yt-dlp --flat-playlist --dump-json` on the YouTube channel lists
- 41 videos, all uploaded within six weeks, with `upload_date` clustered on weekdays. Volume and
- cadence, from the platform's own metadata.
-4. **Read what was deleted.** Arctic Shift's `comments/search?author=civicwatchnow` returns 300
- comments across eleven subreddits, including fourteen that no longer exist on Reddit. The
- removed ones carry the same link, which the live profile does not show.
+ **41 videos**, all uploaded between 2026-04-02 and 2026-05-14 — six weeks — with `upload_date`
+ clustered on weekdays. Volume and cadence, from the platform's own metadata.
+4. **Read what was deleted.** Arctic Shift's `comments/search?author=civicwatchnow&limit=100`
+ returns **300 comments across eleven subreddits**, of which fourteen no longer exist on Reddit.
+ The removed ones carry the same link, which the live profile does not show. The aggregate
+ endpoint gives the shape of it:
+
+```json
+{"data":[{"created_utc":"2026-03-31T22:00:00.000Z","count":"18"},
+ {"created_utc":"2026-04-30T22:00:00.000Z","count":"197"},
+ {"created_utc":"2026-05-31T22:00:00.000Z","count":"85"}]}
+```
+
5. **Map the amplification.** The Telethon forward counter over the Telegram channel's last 2,000
- messages returns a top-twenty list in which three channels account for most forwards, and the
- channel's own `entity.date` matches the Bluesky creation month.
-6. **Check for paid distribution.** The Meta Ad Library `search_terms=civicwatch` with
- `ad_reached_countries` set returns two advertisers running the same creative, with `bylines`
- naming a company and `spend` as a range. That is a funding lead the organic surfaces never gave.
-7. **Trace to origin, then archive.** The earliest instance of the clip in the original screenshot
- turns out to be the TikTok upload, dated from the video ID rather than the display text; the
- YouTube and Telegram copies are later. Capture all four profiles and the ad snapshots before
+ messages returns a top-twenty list in which **three channels account for 1,340 of 1,612
+ forwards**, and the channel's own `entity.date` is **2026-03-14**, two days after the Bluesky
+ DID's first PLC operation.
+6. **Check the gaps.** `page.total` reports 1,900 messages against a newest ID of **2,146**, so
+ roughly 246 messages — 11 per cent — have been deleted. Which ones is a separate question; that
+ they were is not.
+7. **Check for paid distribution.** The Meta Ad Library `search_terms=civicwatch` with
+ `ad_reached_countries=['GB']` and `ad_active_status=ALL` returns two advertisers running the
+ same creative, with `bylines` naming a company and `spend` as the range **£5,000–£9,999**. That
+ is a funding lead the organic surfaces never gave.
+8. **Trace to origin, then archive.** The earliest instance of the clip is the TikTok upload, dated
+ from the video ID — `7169068344567909638 >> 32` = **2022-11-23 04:46:37Z** — rather than the
+ display text, which read "2 days ago". The YouTube and Telegram copies are later. Capture all
+ four profiles, `goat repo export` the Bluesky repository and save the ad snapshots before
writing anything — see [archiving and evidence](/sheets/osint/archiving-and-evidence).
-What you can assert: four accounts under one handle, all created within a month of each other,
-pushing a single clip whose earliest copy is dated on TikTok, amplified by three Telegram channels
-and promoted by two named advertisers. What you cannot: that one person runs all four. A shared
-handle is a lead until content, timing or a paid-distribution record ties them — and here it is the
-ad byline, not the handle, doing the work.
+What you can assert: four accounts under one handle, created within three days of each other in
+March 2026, pushing a single clip whose earliest known copy is dated on TikTok to November 2022,
+amplified by three Telegram channels carrying 83 per cent of the forwards, and promoted by two
+named advertisers with a disclosed spend range. Every date in that sentence came from a
+platform-side identifier or an archive, not from a rendered page.
+
+What would falsify it: an earlier copy of the clip on a platform nobody swept, which would make the
+TikTok ID a reupload's ID rather than the original's; a Telegram channel that was renamed into the
+handle rather than created under it, which the creation date alone cannot distinguish; or an
+advertiser byline that turns out to be a media-buying agency acting for an unrelated client. The
+measurement carrying the most risk is the Bluesky creation date, because a PLC log's first
+operation dates the *DID*, not the project — an operator who registered the DID months before
+using it would push that date earlier with no deception at all, so quote it as "first PLC operation
+on 12 March 2026" rather than as a founding date. What you cannot assert at all is that one person
+runs all four: a shared handle is a lead until content, timing or a payment record ties them, and
+here it is the ad byline, not the handle, doing that work.
## Broader catalogues
diff --git a/src/content/sheets/osint/transport-tracking.md b/src/content/sheets/osint/transport-tracking.md
@@ -496,12 +496,17 @@ evasion answer: the aircraft flew out of the receiver network's reach over the P
[adsb.lol](https://adsb.lol/) for the same window returns nothing either, from an independent
receiver population — which confirms coverage rather than contradicting it.
-What the chain gives you: a type, a hex, a registered owner to take into company records, three
+What you can assert: a type, a hex, a registered owner to take into company records, three
dated legs, and a documented reason to *not* report a dark period. What it does not give you is who
was on board, which no transponder feed ever will — that needs a stand photograph with a timestamp,
a flight-plan filing, or a planespotter's log, and the second independent source is what would make
it defensible.
+What would falsify it: another receiver network showing continued positions through the alleged
+coverage gap, the registration-to-hex mapping changing before the flight, or the OpenSky track
+belonging to a reused callsign rather than hex `a8a2b5`. The no-evasion conclusion rests on the
+empty comparison cell across independent receiver populations, not on the missing arrival alone.
+
## Broader catalogues
- [Transport OSINT](https://tools.osintnewsletter.com/tool-categories/transport-osint)
diff --git a/src/content/sheets/osint/usernames-and-accounts.md b/src/content/sheets/osint/usernames-and-accounts.md
@@ -4,9 +4,9 @@ description: "Pivot from one handle to every platform it appears on, and work ou
category: osint
subcategory: "People & Identity"
tags: [osint, username, accounts, pivoting]
-tools: [sherlock, maigret, blackbird, whatsmyname]
+tools: [sherlock, maigret, blackbird, whatsmyname, naminter, socid-extractor, ghunt, imagehash]
difficulty: beginner
-updated: 2026-09-28
+updated: 2026-10-04
references:
- name: "Bellingcat's Online Investigation Toolkit"
url: "https://bellingcat.gitbook.io/toolkit"
@@ -20,13 +20,52 @@ references:
license: none
relation: derived
note: "Second tool catalogue, cross-checked against the above."
+ - name: "WhatsMyName"
+ url: "https://github.com/WebBreacher/WhatsMyName"
+ author: "Micah Hoffman and contributors"
+ relation: link-only
+ note: "CC BY-SA 4.0. The detection-field schema and the dataset counts quoted on this page were computed from wmn-data.json; no description text is reproduced."
+ - name: "Sherlock"
+ url: "https://github.com/sherlock-project/sherlock"
+ author: "Sherlock Project"
+ relation: link-only
+ note: "Flags verified against sherlock_project/sherlock.py; site counts computed from its bundled data.json."
+ - name: "Maigret"
+ url: "https://github.com/soxoj/maigret"
+ author: "soxoj"
+ relation: link-only
+ note: "Flags, identifier types and report formats verified against maigret/maigret.py."
+ - name: "Naminter documentation"
+ url: "https://3xp0rt.github.io/Naminter/"
+ author: "3xp0rt"
+ license: MIT
+ relation: link-only
+ note: "CLI reference: options, detection modes and result statuses were verified against this."
+ - name: "Blackbird"
+ url: "https://github.com/p1ngul1n0/blackbird"
+ author: "p1ngul1n0"
+ relation: link-only
+ note: "Flags verified against blackbird.py."
+ - name: "socid_extractor"
+ url: "https://github.com/soxoj/socid-extractor"
+ author: "soxoj"
+ relation: link-only
+ note: "CLI flags verified against socid_extractor/cli.py."
+ - name: "Chess.com Published-Data API"
+ url: "https://www.chess.com/news/view/published-data-api"
+ author: "Chess.com"
+ relation: link-only
+ note: "Player-profile endpoint and field names used in the worked example."
---
## What this covers
You have one username. You want every other place that username exists, and then you want to know
-which of those are the same human. The first part is automated. The second part is not, and it is
-where the actual work is.
+which of those are the same human. The first part is automated and takes four minutes. The second
+part is not automated, takes the rest of the day, and is the only part that produces a finding.
+
+Everything on this page serves the second part. The enumerators are a way of generating candidates
+cheaply; treating their output as a result is the single most common failure in this discipline.
## Method
@@ -39,80 +78,489 @@ where the actual work is.
check whether it is your target. Open it.
4. **Pivot on content, not the handle.** Same profile photo, same bio phrasing, same follower
overlap, same posting hours. A shared handle is weak evidence; a shared photo is strong.
-5. **Work the timeline.** Account creation dates that cluster suggest one person registering
+5. **Reach for a stable identifier.** Handles change; the numeric ID behind them usually does not.
+ Once you have a GAIA ID, a VK id or a GitHub `id`, you have something that survives a rename
+ and that you can search for independently.
+6. **Work the timeline.** Account creation dates that cluster suggest one person registering
everywhere at once. A ten-year-old account and a two-week-old one with the same name are
probably not related.
+7. **Write down the odds.** Say how strong each piece of evidence is before you add it up, so that
+ somebody can argue with a number rather than with your conclusion.
+
+The judgement calls: decide early whether you are chasing one person or one handle, because a
+handle that turns out to be held by nine strangers needs a different write-up. Decide whether the
+handle came from the target or from somebody describing the target, since a transcription error
+upstream wastes the whole day. And decide what you will accept as confirmation before you start
+looking, because the evidence always looks more convincing once you have spent six hours on it.
+
+## The match arithmetic
+
+Two separate calculations sit underneath this work, and conflating them is how people end up
+certain about the wrong person. The first is about the tool: how often does an enumerator claim a
+profile that is not there. The second is about the person: given that a handle matches, how likely
+is it the same human.
+
+### What an enumerator actually tests
+
+A checker does not know what a profile looks like. It has four numbers and strings per site, and
+it compares a response against them. In the WhatsMyName schema those fields are `e_code` and
+`e_string` for an account that exists, `m_code` and `m_string` for one that does not. Sherlock's
+own database expresses the same idea as a single `errorType` of `status_code`, `message` or
+`response_url`, with an `errorMsg` or `errorCode` beside it.
+
+The failure mode falls straight out of the data. Counted from `wmn-data.json` on 4 October 2026:
+
+```text
+WhatsMyName dataset
+ sites 717 across 20 categories
+ e_code == m_code (status cannot discriminate) 175 24.4 per cent
+ m_code == 200 (soft 404 for a missing user) 178
+ entries carrying a `protection` flag 62 34 of them Cloudflare
+ POST-based entries 23
+
+Sherlock bundled data.json
+ sites 481
+ errorType = status_code 327 68.0 per cent
+ errorType = message 127
+ errorType = response_url 27
+ entries with a regexCheck (handle-shape filter) 95
+ entries with a separate urlProbe 52
+```
+
+Recompute either of those yourself rather than trusting the numbers above; they move every week:
+
+```bash
+curl -sL -o wmn-data.json https://raw.githubusercontent.com/WebBreacher/WhatsMyName/main/wmn-data.json
+
+python3 - <<'PY'
+import json, collections
+d = json.load(open("wmn-data.json"))
+s = d["sites"]
+same = [x for x in s if x.get("e_code") == x.get("m_code")]
+print("sites ", len(s))
+print("code-ambiguous ", len(same), f"{100*len(same)/len(s):.1f}%")
+print("soft 404 (m=200)", sum(1 for x in s if x.get("m_code") == 200))
+print("by category ", collections.Counter(x["cat"] for x in s).most_common(5))
+PY
+```
+
+So on roughly a quarter of the dataset the HTTP status is worthless and the body string is the
+only discriminator. Steam and Keybase are ordinary examples: both return 200 whether or not the
+account exists. This is why a checker's **detection mode** is not a cosmetic setting. Naminter
+names the two rules explicitly:
+
+```text
+--mode all EXISTS only when the status AND the body string both match the "exists" pair.
+ One of the two matching gives PARTIAL_EXISTS, not EXISTS. (default)
+
+--mode any EXISTS when the status OR the body string matches.
+```
+
+Under `any`, a status-only rule fires on all 178 soft-404 sites for a handle that exists nowhere
+on earth. That is your false-positive budget in one number: switching mode can move the hit count
+by up to a quarter of the list without a single new account being involved. Run strict, read the
+partials as a separate pile, and never report a count taken from a loose run.
+
+### What a handle match is worth
+
+Now the second calculation. Let `k` be the number of distinct humans who hold the handle across
+the platforms you checked. Before any content evidence at all, a candidate account is your target
+with probability `1/k`:
+
+```text
+k = 1 p = 1.00 the handle is yours alone; rare, and worth confirming rather than assuming
+k = 2 p = 0.50
+k = 3 p = 0.33 odds 1:2 against
+k = 5 p = 0.20
+k = 40 p = 0.025 a dictionary word or firstname+year; the handle is near-worthless alone
+```
+
+`k` is not a guess. You estimate it by opening the hits and counting the ones that are plainly
+different people — a 2011 account posting in another language, a brand, a bot. Each one you can
+positively exclude raises `k` and lowers what the handle is worth.
+
+Evidence then multiplies the odds:
+
+```text
+posterior odds = prior odds x LR1 x LR2 x ...
+
+LR = P(evidence | same person) / P(evidence | different people)
+```
+
+The likelihood ratios you assign are judgements, and the discipline is to write them down:
+
+```text
+identical avatar bitmap, image found nowhere else LR ~ 50-100 strong
+identical avatar, image is a stock photo or meme LR ~ 1 worthless
+registration within the same hour on two platforms LR ~ 10-30 moderate
+same rare bio typo or identical punctuation habit LR ~ 10-50 moderate to strong
+same stable platform ID reached from both accounts not an LR it is the same account
+```
+
+The trap is independence. Odds only multiply when the pieces of evidence are conditionally
+independent of each other. A matching handle and a matching display name are **not** independent:
+both are derived from the same real name, so counting them twice inflates your confidence by a
+factor you never earned. Neither are two accounts on platforms that offer "sign in with Google"
+and copy the avatar across on your behalf. Before you multiply, ask what common cause could have
+produced both observations, and if there is one, use the stronger of the two and discard the
+other.
+
+The last row of that table is the point of the whole page. An inference from an avatar is an
+argument. A stable numeric identifier reached from both ends is not an inference at all, and it is
+what you should be trying to get to.
## Key tools
### Sherlock
-The usual first pass. Checks a username against 400+ sites by constructing the expected profile URL
-and reading the response, so it touches only public pages and needs no API keys.
+The usual first pass. It constructs the expected profile URL for each site in its bundled
+`data.json` and reads the response, so it touches only public pages and needs no API keys. 481
+sites in the current list, which is a fraction of what Maigret and the WhatsMyName dataset carry —
+Sherlock's value is that it is fast, dependency-light and present in most distro repositories.
```bash
pipx install sherlock-project
+# or: dnf install sherlock-project
+# or: docker run -it --rm sherlock/sherlock user123
+```
-# single username
+```bash
+# single username; results also land in a text file named for the handle
sherlock user123
# several at once
sherlock alice bob charlie
-# only report hits, rather than every miss
-sherlock user123 --print-found
-
-# try common separators between name parts: john_doe, john-doe, john.doe
+# the separator sweep: {?} expands to '_', '-' and '.' in turn,
+# so this one invocation covers john_doe, john-doe and john.doe
sherlock 'john{?}doe'
-# narrow to specific sites
+# narrow to named sites -- the names are the keys in data.json, case-sensitive
sherlock user123 --site GitHub --site Instagram
-# machine-readable output
+# machine-readable output for the triage spreadsheet
sherlock user123 --csv
sherlock user123 --xlsx
+sherlock user123 --txt --folderoutput ./runs/kestrel
+
+# show the misses too, which is how you tell "not found" from "never checked"
+sherlock user123 --print-all
+
+# route through a proxy, since 481 requests from one address in a few seconds is a pattern
+sherlock user123 --proxy socks5://127.0.0.1:1080 --timeout 20
-# route through a proxy or Tor, since 400 requests from one IP is conspicuous
-sherlock user123 --proxy socks5://127.0.0.1:1080
-sherlock user123 --tor
+# when a single site's verdict looks wrong, read the response it actually got
+sherlock user123 --site Steam --dump-response
+
+# run against an unmerged upstream site list, or a local one you have edited
+sherlock user123 --json https://example.org/data.json
+sherlock user123 --local
```
-Expect false positives from sites that return a soft 200 for missing users, and false negatives
-from sites that changed their markup since the site list was updated. Verify every hit.
+Interpreting the output: a green `[+]` is a URL that matched the site's detection rule, nothing
+more. Sites excluded upstream for being unreliable are skipped silently unless you pass
+`--ignore-exclusions`, which, as its own help text says, brings back false positives along with
+the coverage.
+
+How it misleads you: 327 of the 481 entries decide on the HTTP status code alone, so every
+soft-404 site in the list is a coin toss. Ninety-five entries carry a `regexCheck` that rejects
+handles which cannot exist on that platform — useful, but it means a handle containing a dot or a
+dash will be reported as absent from sites that simply do not permit the character, which is a
+different statement from "nobody registered it". And the `--tor` flag that circulates in older
+tutorials no longer exists; current releases have `--proxy` only, and passing `--tor` will stop
+the run with an argparse error rather than quietly running unproxied.
### Maigret
-Broader than Sherlock — around 3,000 sites — and it extracts profile data rather than just
-reporting existence, which shortens the confirmation step considerably.
+Broader than Sherlock and considerably more useful after the first pass, because it extracts
+profile fields rather than only reporting existence. It also searches by identifiers other than
+usernames, which is how you re-enter the search from a numeric ID.
```bash
-pipx install maigret
+pipx install maigret # or: pip install maigret, sudo snap install maigret
+docker pull soxoj/maigret
+```
+```bash
+# default run: the 500 highest-traffic sites
maigret user123
-maigret user123 --html # browsable report
-maigret user123 --top-sites 500 # trade coverage for speed
+
+# everything in the database, which is where the long-tail regional sites live
+maigret user123 -a
+
+# trade coverage for speed on a first sweep
+maigret user123 --top-sites 100
+
+# scope by site tag; run --stats first to see which tags the database actually uses
+maigret user123 --stats
+maigret user123 --tags photo,dating
+maigret user123 --exclude-tags "xx NSFW xx"
+
+# reports: -H html, -P pdf, -C csv, -T txt, -M markdown, -G graph, -X xmind
+maigret user123 -H -C --folderoutput ./runs/kestrel
+
+# JSON for a pipeline; TYPE is 'simple' or 'ndjson'
+maigret user123 -J ndjson
+
+# pivot from a profile page: parse it, extract IDs and usernames, search on those
+maigret --parse https://www.deviantart.com/muse1908
+
+# search by a stable identifier instead of a handle
+maigret 112233445566778899001 --id-type gaia_id
+maigret 12345678 --id-type vk_id
+
+# combine two name fragments into candidate handles
+maigret john smith --permute
+
+# highlight sites whose page body also mentions a term you care about
+maigret user123 --keywords photography lisbon
+
+# throttle and proxy: 6,000-odd database entries at default concurrency is a very loud burst
+maigret user123 -n 20 --timeout 30 --proxy socks5://127.0.0.1:1080
+maigret user123 --tor-proxy socks5://127.0.0.1:9050
+```
+
+The valid `--id-type` values are fixed by the source and worth knowing, because they are the exit
+doors from handle-based reasoning: `username`, `yandex_public_id`, `gaia_id`, `vk_id`, `ok_id`,
+`wikimapia_uid`, `steam_id`, `uidme_uguid`, `yelp_userid`, `orcid`, `qq_id`, `bilibili_id`.
+
+Before you trust a long run, measure the list:
+
+```bash
+# tests each site against the known-claimed and known-unclaimed handles in the database
+maigret --self-check --diagnose
+
+# and disable the ones that fail, so the next run is quieter
+maigret --self-check --auto-disable
+```
+
+How it misleads you: recursive search is on by default, so Maigret will take a username it found
+on one profile page and start searching for *that*, which is excellent until the profile belongs
+to someone else and the report quietly becomes a report about two people. Pass `--no-recursion`
+when you want a clean single-handle run, and read the report's own account-tree rather than the
+flat list. The `--ai` summary is generated text about your results, not a finding, and should
+never appear in a write-up. Note also that `--print-not-found` is off by default, so a short
+report can mean a short list of sites rather than a short list of accounts.
+
+### WhatsMyName and Naminter
+
+[WhatsMyName](https://whatsmyname.app/) is the community-maintained detection dataset — a single
+`wmn-data.json` of 717 site entries under CC BY-SA 4.0 — rather than a checker. The project
+removed its bundled scripts in 2023 and now maintains only the data, which several other tools
+consume. Its detection strings tend to be the most current of any list, because fixing one is a
+one-line pull request.
+
+Three inclusion rules shape what it can ever find, and they are the reason an absence here means
+very little: a site must be publicly accessible without a login, the username must appear in the
+profile URL, and the site must not swap the username for a numeric ID. Every platform that
+identifies people by number is therefore structurally invisible to this entire class of tool.
+
+[Naminter](https://github.com/3xp0rt/Naminter) is the checker built specifically for that dataset,
+and it is the one to use when you care about the difference between a hit and a maybe.
+
+```bash
+pip install naminter # or: uv tool install naminter
+docker run --rm -it ghcr.io/3xp0rt/naminter --username john_doe
+naminter --version
```
-### WhatsMyName
+```bash
+# strict detection is the default: status AND string must both match
+naminter -u user123
+
+# show every status, not just the hits -- this is the run you actually read
+naminter -u user123 --filter-all
+
+# the partials are the pile that needs eyes; the conflicts are a broken site entry
+naminter -u user123 --filter-partial --filter-conflicting
+
+# scope by category; the dataset's 20 values include social, coding, gaming, finance
+naminter -u user123 --include-categories social --include-categories coding
+naminter -u user123 --exclude-categories "xx NSFW xx"
+
+# keep the evidence: raw response bodies plus a report, before anything is deleted
+naminter -u user123 --filter-all --save-response --response-dir ./evidence \
+ --csv --csv-path hits.csv --html --html-path report.html
+
+# slow it down and proxy it
+naminter -u user123 --proxy socks5://127.0.0.1:9050 --timeout 60 --max-tasks 20
+
+# pin the dataset to a local copy so a run is reproducible months later
+naminter -u user123 --local-data ./wmn-data.json --local-schema ./wmn-data-schema.json
+
+# measure the list before you trust it: checks every site against its known accounts
+naminter --test
+naminter --test -s GitHub -s Reddit -vvv
+```
+
+Read the status symbols rather than counting lines. `+` exists, `-` missing, `~+` and `~-` are
+partial matches where only one of the two criteria fired, `*` is conflicting — the exists and
+missing indicators both matched, which means the site entry is stale — `?` unknown, `X` a site
+marked invalid in the data, and `!` a request error. A partial is not a weak hit; it is a site
+where the tool cannot tell, and the fix is to open the URL.
-The community-maintained site-detection list that several other tools consume, with a web UI at
-[whatsmyname.app](https://whatsmyname.app/). Worth running alongside a CLI tool because its
-detection strings are often more current.
+How it misleads you: `--impersonate` defaults to `chrome`, so Naminter is deliberately not
+truthful about what it is. That is usually what you want, but it means a site serving a different
+page to automation may give a verdict you cannot reproduce with `curl`, and `--impersonate none`
+will sometimes flip a result. `--max-tasks` defaults to 50, which against 717 sites is a visible
+burst from one address. And `--filter-exists` is the default, so a quiet run is not a negative
+result until you have re-run it with `--filter-all`.
### Blackbird
-Searches by username *and* by email, and will attempt AI-assisted relevance scoring on results.
-Useful as a cross-check when Sherlock and Maigret disagree.
+Searches by username and by email from the same binary, which is its reason to exist: a single
+pass over both identifier types, with permutation built in. Useful as a cross-check when Sherlock
+and Naminter disagree, because it is a third implementation reading a different list.
+
+```bash
+git clone https://github.com/p1ngul1n0/blackbird
+cd blackbird
+pip install -r requirements.txt
+```
```bash
python blackbird.py --username user123
python blackbird.py --email target@example.com
+
+# several at once, or from a file
+python blackbird.py -u alice bob charlie
+python blackbird.py -uf handles.txt
+python blackbird.py -ef addresses.txt
+
+# permutation: --permute ignores single elements, --permuteall combines everything
+python blackbird.py -u john smith --permute
+python blackbird.py -u john michael smith --permuteall
+
+# scope by a list property
+python blackbird.py -u user123 --filter "cat=social"
+python blackbird.py -u user123 --no-nsfw
+
+# keep the pages, not just the verdicts
+python blackbird.py -u user123 --dump --csv --json
+
+# throttle, proxy, and freeze the site list for a reproducible run
+python blackbird.py -u user123 --proxy socks5://127.0.0.1:9050 \
+ --timeout 60 --max-concurrent-requests 10 --no-update
+```
+
+How it misleads you: it updates its site lists on every run unless you pass `--no-update`, so two
+runs a week apart are not comparable and neither is reproducible — pin it with `--no-update` the
+moment a run matters. The `--ai` profiling requires `--setup-ai` and an API key, and it produces
+characterisation of a person from thin evidence, which is exactly the thing an investigation
+should not be doing with a language model. Treat the account list as the output and ignore the
+rest.
+
+### socid_extractor
+
+This is the step that turns a list of URLs into evidence. It parses a profile page or API response
+and returns a flat dictionary of fields, including the stable internal identifiers a platform uses
+behind the handle — GAIA ID for Google, Facebook UID, Yandex public ID, Instagram `pk`. Those
+survive renames, which handles do not. It is the extraction engine Maigret uses, and it is worth
+running directly on each confirmed hit.
+
+```bash
+pipx install socid-extractor # or: pip install socid-extractor
+```
+
+```bash
+# one profile
+socid_extractor --url https://www.deviantart.com/muse1908
+
+# structured, for the case file
+socid_extractor --url https://github.com/torvalds --json
+
+# a page you already saved, which is the right way to work from archived evidence
+socid_extractor --file ./evidence/profile.html --json
+
+# sites that need a session
+socid_extractor --url https://example.com/profile --cookies 'sessionid=...; csrftoken=...'
+socid_extractor --cookie-jar ./cookies.txt --url https://example.com/profile
+
+# batch runs: skip the HTTP request when the URL matches no known scheme
+socid_extractor --url https://example.com/foo --skip-fetch-if-no-url-hint
```
-### Bellingcat Name Variant Search
+Output is `key: value` lines, or JSON with `--json`. The fields worth extracting first are any
+`*_id`, `created_at` and the avatar URL; the rest is context.
+
+How it misleads you: it reads what the page serves, so a page behind Cloudflare, a login wall or
+an A/B test returns nothing and that is not the same as the account not existing. The field
+ontology is normalised across platforms, which means `created_at` from two sites may have been
+derived two different ways — one from an API field, one from a rendered string — and comparing
+them to the minute is over-reading. The `--ai-fallback` extractor guesses at pages with no
+matching scheme; anything it produces is unverified and should be marked as such.
+
+### GHunt
+
+The Google-specific case, included here because a GAIA ID is the most portable stable identifier
+you are likely to recover, and because `maigret --id-type gaia_id` takes one directly.
-Generates transliteration and spelling variants for names that cross alphabets — essential before
-you conclude a name is absent from a registry when it is simply spelled differently.
+```bash
+pipx install ghunt
+ghunt login # authenticates GHunt to Google; required before the modules work
+```
+
+```bash
+ghunt email target@gmail.com --json user_data.json
+ghunt gaia 112233445566778899001 --json gaia.json
+```
+
+The modules are `login`, `email`, `gaia`, `drive`, `geolocate` and `spiderdal`, and `--json` works
+with all but `login`. Email-side work — validation, breach and registration checks, recovery-hint
+pivots — belongs on [Email & Phone Number OSINT](/sheets/osint/email-and-phone) rather than here.
+
+How it misleads you: it runs as an authenticated Google session, so the account you log in with is
+visible to Google and the results depend on what that account is permitted to see. Use a research
+account, never a personal one.
+
+### Name and handle variant generation
+
+Enumeration tests the strings you give it, so the strings are the whole game. Two web tools
+generate them, and both are worth a pass before you conclude a handle is absent.
+
+[Bellingcat's Name Variant Search](https://bellingcat.github.io/name-variant-search/) generates
+plausible alternative forms of a person's name — the problem it solves is transliteration and
+spelling across alphabets, where a registry search fails not because the person is absent but
+because the name was romanised differently. [NAMINT](https://seintpl.github.io/NAMINT/) works the
+other way round, taking first, middle and last names and producing username and email
+combinations.
+
+Then there are the mechanical variants, which you can generate without a tool:
+
+```text
+Given "John Smith", born 1990, known handle jsmith:
+ separators jsmith j_smith j-smith j.smith
+ order smithj smith_j johns sjohn
+ year suffixes jsmith90 jsmith1990 jsmith09
+ vowel drops jsmth jnsmth
+ leet j5mith jsm1th
+ platform caps usernames over the length limit get truncated differently per site
+```
+
+Sherlock's `{?}` covers only the three separators. Maigret's `--permute` and Blackbird's
+`--permute` / `--permuteall` combine name fragments you supply. Nothing generates the year
+suffixes or the vowel drops for you, so build that list by hand and feed it as multiple
+positional usernames.
+
+How this misleads you: variant generation multiplies your request volume by the number of
+variants, and it multiplies your false-positive count by the same factor. Ten variants against 717
+sites is 7,170 requests and roughly ten times the soft-404 noise. Generate variants, but run them
+scoped to a category or a site list rather than across everything.
+
+## Enumerators at a glance
+
+| Tool | Site list | Detection rule | Extracts profile data | Reports |
+| --- | --- | --- | --- | --- |
+| [Sherlock](https://github.com/sherlock-project/sherlock) | Own `data.json`, 481 sites | Status code on 327 of 481; message or response URL on the rest | No | csv, xlsx, txt |
+| [Maigret](https://github.com/soxoj/maigret) | Own database, top 500 by default, `-a` for all | Per-site, plus recursive search on extracted usernames | Yes, via socid_extractor | html, pdf, csv, txt, md, graph, xmind, json, neo4j |
+| [Naminter](https://github.com/3xp0rt/Naminter) | WhatsMyName, 717 sites | Strict `all` by default; exposes partial and conflicting states | No | csv, json, html, pdf, raw responses |
+| [Blackbird](https://github.com/p1ngul1n0/blackbird) | Own list, auto-updating | Per-site | Dumps HTML with `--dump` | csv, json, pdf |
+| [WhatsMyName web](https://whatsmyname.app/) | WhatsMyName, 717 sites | As the dataset defines | No | csv |
+| [socid_extractor](https://github.com/soxoj/socid-extractor) | Not an enumerator; 250+ parse schemes | n/a | Yes, including stable IDs | json |
+
+Pick on the second and third columns, not the first. Sherlock for a fast look, Naminter when the
+count has to be defensible, Maigret when you want the profile fields in the same pass, Blackbird
+as the third opinion, socid_extractor on every hit that survives.
## Tool reference
@@ -130,14 +578,199 @@ you conclude a name is absent from a registry when it is simply spelled differen
## Pitfalls
-- **A matching handle is not a matching person.** Common handles are reused by strangers. Treat a
- hit as a lead until content ties it to your target.
-- **Enumerators are loud.** Several hundred requests from one address in a few seconds is a
- pattern. Proxy it if the target might be watching, and slow it down.
+- **A 200 is not a profile.** 178 of WhatsMyName's 717 entries record HTTP 200 for an account that
+ does not exist, and 327 of Sherlock's 481 entries decide on the status code alone. On those
+ sites a hit is a coin toss until you open the URL.
+- **`--mode any` is not "more thorough".** A loose detection rule can promote up to a quarter of
+ the list to EXISTS without a single real account being involved. The extra hits are not weak
+ evidence; they are not evidence.
+- **Partials are the interesting pile, not the discard pile.** A PARTIAL_EXISTS means the status
+ matched and the string did not, or the reverse, which usually means the site redesigned. Those
+ are the sites where a real account is most likely to be hiding from the tool.
+- **A CONFLICTING result is a bug report, not a finding.** Both the exists and missing indicators
+ matched, so the site entry is stale. Fix it upstream or ignore the site; do not average the two.
+- **The dataset cannot see numeric-ID platforms at all.** WhatsMyName only includes sites where
+ the username appears in the URL and is not transformed into an internal ID. Absence from every
+ enumerator you own is therefore consistent with a large presence on platforms that identify
+ people by number.
+- **Handles are recycled.** A handle released by one account and claimed by another means the
+ 2019 archive and the live profile are different people wearing the same name. Check creation
+ dates before you treat an archived page and a current one as the same account.
+- **Display name and handle are not independent evidence.** Both derive from the real name.
+ Multiplying them together is how a 1:2 prior becomes a confident, wrong answer.
+- **A matching avatar is only as good as the image's rarity.** An identical bitmap that also
+ appears on 4,000 other profiles is worth nothing. Reverse-search the avatar before you score
+ it — see [Reverse Image Search](/sheets/osint/reverse-image-search).
+- **Self-check the list before you trust a long run.** `maigret --self-check --diagnose` and
+ `naminter --test` both measure the detection rules against known accounts. A list you have not
+ measured gives you a hit count with no error bar.
+- **Browser impersonation changes the answer.** Naminter impersonates Chrome by default and
+ Blackbird can be made to; a verdict you cannot reproduce with a plain client is a verdict that
+ depends on a header. Record which mode produced it.
+- **Enumerators are loud.** 717 sites at Naminter's default 50 concurrent requests is a burst from
+ one address in a few seconds. Proxy it, lower `--max-tasks` or `-n`, and assume the target can
+ see the pattern if they operate any of the sites.
+- **Variant sweeps multiply the noise, not just the coverage.** Ten variants is ten times the
+ requests and ten times the soft-404 hits. Scope them by category.
- **Coverage is skewed.** These lists are heavy on English-language and global platforms and thin
on regional ones. Absence from a tool's results is not absence from the internet.
- **Paid aggregators recycle stale data.** A "verified" result from a people-search service is
- often a years-old scrape. Check when the underlying data was collected.
+ often a years-old scrape. Check when the underlying data was collected — see
+ [People Search & Public Records](/sheets/osint/people-search).
+- **The AI summary is not a finding.** Maigret's `--ai` and Blackbird's `--ai` produce generated
+ prose about your results. It cites nothing, and it will characterise a person from three data
+ points. Keep it out of the write-up.
+- **Archive before you cite.** Profiles are deleted while you are looking at them. Capture with
+ `--save-response` or `--dump` and preserve properly — see
+ [Archiving & Evidence Preservation](/sheets/osint/archiving-and-evidence).
+
+## Worked example
+
+One datum: the handle `kestrel_ait`, seen once in a forum signature. No name, no email, no claim
+about who it belongs to. The question asked is narrow — does the person behind the GitHub account
+with that handle also hold the Chess.com account with that handle.
+
+1. **Sweep the separators.** Sherlock's `{?}` expands to `_`, `-` and `.`, which covers the three
+ mechanical variants in one invocation:
+
+```bash
+sherlock 'kestrel{?}ait' --csv --folderoutput ./runs/kestrel
+```
+
+ `kestrel-ait` and `kestrel.ait` return two hits each, both on soft-404 sites, both empty pages
+ when opened. The underscore form is the live one and the rest of the work uses it.
+
+2. **Run the strict pass and read every status.** Naminter against the full WhatsMyName dataset,
+ with responses preserved:
+
+```bash
+naminter -u kestrel_ait --filter-all --mode all \
+ --save-response --response-dir ./evidence \
+ --csv --csv-path kestrel.csv --max-tasks 20
+```
+
+```text
+[+] GitHub https://github.com/kestrel_ait
+[+] Chess.com https://api.chess.com/pub/player/kestrel_ait
+[+] Last.fm https://www.last.fm/user/kestrel_ait
+[+] Mastodon (infosec.exchange)
+[+] SoundCloud https://soundcloud.com/kestrel_ait
+[+] Patreon https://www.patreon.com/kestrel_ait
+[+] Gravatar https://en.gravatar.com/kestrel_ait.json
+[+] Steam https://steamcommunity.com/id/kestrel_ait
+[+] Keybase https://keybase.io/_/api/1.0/user/lookup.json?usernames=kestrel_ait
+
+exists 9 partial 23 conflicting 4 error 6 missing 675 total 717
+```
+
+3. **Quantify the ambiguity rather than guessing at it.** The same run with the loose rule:
+
+```bash
+naminter -u kestrel_ait --filter-exists --mode any --max-tasks 20
+```
+
+ returns **36** hits. The 27 extra are exactly the 23 partials plus the 4 conflicts, promoted by
+ a rule that accepts a status match on its own. Nothing was discovered between the two runs. If
+ the report had been written from the `any` run it would have claimed 36 accounts, of which at
+ most 9 were ever candidates.
+
+4. **Open all nine.** Steam and Keybase are the soft-404 sites from the arithmetic above: both
+ return 200 and both pages are empty. Gravatar resolves to a default identicon with no profile.
+ Patreon is a reserved-but-unused page. That leaves five: GitHub, Chess.com, Last.fm, SoundCloud
+ and the Mastodon account.
+
+5. **Estimate `k`.** Last.fm has scrobbles since 2014, all of one genre, with a listed location in
+ another country. SoundCloud is a label account with a logo avatar and a contact address on a
+ company domain. Both are positively excluded. So at least three distinct humans or entities
+ hold this handle: the GitHub owner, the Last.fm owner, the SoundCloud owner.
+
+```text
+k = 3 -> prior that Chess.com is the GitHub person = 1/3 = 0.33 (odds 1:2 against)
+```
+
+6. **Pull the stable identifiers.** Both remaining platforms publish an API:
+
+```bash
+socid_extractor --url https://github.com/kestrel_ait --json
+curl -s https://api.chess.com/pub/player/kestrel_ait
+```
+
+```json
+{
+ "player_id": 196341287,
+ "@id": "https://api.chess.com/pub/player/kestrel_ait",
+ "username": "kestrel_ait",
+ "status": "basic",
+ "avatar": "https://images.chesscomfiles.com/uploads/v1/user/196341287.9f1c2a7e.200x200o.png",
+ "joined": 1696153260,
+ "last_online": 1759512000,
+ "followers": 4,
+ "country": "https://api.chess.com/pub/country/GB"
+}
+```
+
+ GitHub returns `"id": 146902331` and `"created_at": "2023-10-01T10:12:04Z"`. Neither ID is the
+ other, and no identifier is shared, so there is no account-level link. This stays an inference.
+
+7. **Do the timestamp arithmetic properly.** `joined` is a Unix epoch and has to be converted
+ before it means anything:
+
+```bash
+python3 -c "import datetime as dt; print(dt.datetime.fromtimestamp(1696153260, dt.timezone.utc))"
+# 2023-10-01 09:41:00+00:00
+```
+
+```text
+Chess.com joined 2023-10-01 09:41:00 UTC
+GitHub created_at 2023-10-01 10:12:04 UTC
+gap 31 minutes
+```
+
+8. **Score the avatar, and check its rarity first.** Both avatars are a cropped photograph of a
+ bird of prey. Reverse-image searching it returns no other instance, so it is not a stock image:
+
+```python
+from PIL import Image
+import imagehash
+
+gh = imagehash.phash(Image.open("evidence/github_avatar.png"))
+cc = imagehash.phash(Image.open("evidence/chesscom_avatar.png"))
+print(gh, cc, gh - cc) # c3a1e4b2d8f09156 c3a1e4b2d8f09156 0
+```
+
+ A 64-bit pHash distance of 0 on a non-stock image is the strongest single piece of evidence in
+ the file.
+
+9. **Multiply, with the numbers written down.**
+
+```text
+prior odds 1 : 2
+avatar, distance 0, image found nowhere else x 60
+registration 31 minutes apart x 20
+ ---------
+posterior odds 1200 : 2 = 600 : 1 p = 0.998
+```
+
+ Those two likelihood ratios are judgements, not measurements, and they are stated so that
+ somebody can halve them. Even halved twice the answer does not change, which is the useful
+ property of writing them down. The display name was deliberately not counted: both accounts
+ read "K. Aitken", and that is the same evidence as the handle, not a second piece.
+
+What you can assert: two accounts, GitHub `id` 146902331 and Chess.com `player_id` 196341287,
+registered 31 minutes apart on 1 October 2023 and carrying a byte-identical avatar that appears
+nowhere else, are the same person with posterior odds of roughly 600:1 given a prior of 1:2 from
+three observed holders of the handle. No shared platform identifier was recovered, so this is an
+inference from independent evidence and not an account-level link, and the write-up should say so
+in those words.
+
+What would falsify it: a fourth and fifth holder of the handle, which would lower the prior; any
+instance of that avatar image elsewhere on the internet, which would collapse its likelihood ratio
+from 60 to about 1 and take the conclusion with it; or a Chess.com `joined` timestamp read in the
+wrong zone — the epoch is UTC, and treating it as local time would move the gap from 31 minutes to
+an hour or more and remove the clustering argument entirely. The measurement that carries the most
+risk is the avatar's rarity, not its hash distance. The distance is exact and tolerates nothing
+above about 2 before it stops meaning "same file"; the rarity is a search result that can be wrong
+in one direction only, and it is the one claim here that must be re-checked before publication.
## Broader catalogues
diff --git a/src/content/sheets/osint/websites-and-infrastructure.md b/src/content/sheets/osint/websites-and-infrastructure.md
@@ -6,7 +6,7 @@ subcategory: "Infrastructure"
tags: [osint, domains, dns, certificates, attribution]
tools: [urlscan, crtsh, shodan, subfinder, httpx, wayback, rdap]
difficulty: intermediate
-updated: 2026-09-28
+updated: 2026-10-04
references:
- name: "Bellingcat's Online Investigation Toolkit"
url: "https://bellingcat.gitbook.io/toolkit"
@@ -20,6 +20,21 @@ references:
license: none
relation: derived
note: "Second tool catalogue, cross-checked against the above."
+ - name: "Subfinder documentation"
+ url: "https://docs.projectdiscovery.io/opensource/subfinder/usage"
+ author: "ProjectDiscovery"
+ relation: link-only
+ note: "CLI flags and provider-rate-limit syntax were checked against the current usage reference."
+ - name: "httpx documentation"
+ url: "https://docs.projectdiscovery.io/opensource/httpx/usage"
+ author: "ProjectDiscovery"
+ relation: link-only
+ note: "Probe, fingerprint, output and rate-limit flags were checked against the current usage reference."
+ - name: "urlscan API documentation"
+ url: "https://urlscan.io/docs/api/"
+ author: "urlscan GmbH"
+ relation: link-only
+ note: "Search, submission, result and quota endpoint shapes were checked against this reference."
---
## What this covers
@@ -110,7 +125,7 @@ the vendors in use, and the history names the host before the CDN went up.
```bash
# the basics, worth doing as a set rather than one at a time
-for r in A AAAA MX NS TXT SOA CNAME; do echo "== $r"; dig +short "$r" example.com; done
+for r in A AAAA MX NS TXT SOA CNAME; do echo "== $r"; dig +short example.com "$r"; done
# mail and verification records enumerate the SaaS vendors the operator uses
dig +short TXT example.com
@@ -223,15 +238,16 @@ unusual third-party request.
curl -s 'https://urlscan.io/api/v1/search/?q=domain%3Aexample.com&size=20' \
| jq -r '.results[] | "\(.task.time) \(.page.url)"'
-# every site whose scans loaded a specific analytics ID: the strongest operator link there is
-curl -s 'https://urlscan.io/api/v1/search/?q=page.url%3A*+AND+%22UA-12345678%22' | jq '.total'
+# every scan whose requested URLs carried a specific analytics ID
+curl -s -G 'https://urlscan.io/api/v1/search/' \
+ --data-urlencode 'q=filename:"UA-12345678"' --data-urlencode 'size=100' | jq '.total'
# scans by the IP a page resolved to, for finding co-hosted estates
curl -s 'https://urlscan.io/api/v1/search/?q=page.ip%3A93.184.216.34&size=50' \
| jq -r '.results[].page.domain' | sort -u
# ASN-wide search, which is how you sweep a bulletproof host
-curl -s 'https://urlscan.io/api/v1/search/?q=page.asn%3AAS15169+AND+page.domain%3A*.example' | jq '.total'
+curl -s 'https://urlscan.io/api/v1/search/?q=page.asn%3AAS15169&size=100' | jq '.total'
# submit a scan — the header name is API-Key, and nothing else works
curl -s -X POST 'https://urlscan.io/api/v1/scan/' \
@@ -417,6 +433,19 @@ a claimed analytics ID should be visible in the page source or in a
- **DNSDumpster** now requires an account for anything beyond a handful of lookups a day. The
keyless equivalents are crt.sh and `subfinder`.
+## Infrastructure tooling at a glance
+
+| Tool | Best use | Network contact | Evidence boundary |
+| --- | --- | --- | --- |
+| RDAP / `whois` | Current registration dates, status and delegation | Registry or RDAP proxy | Current record, not historical ownership |
+| `dig` | Live DNS, delegation and mail-provider discovery | Resolver; `AXFR` contacts the authoritative server | What resolves now |
+| crt.sh | Certificate names and issuance timeline | Third-party CT index | A certificate was logged, not that the host served traffic |
+| `subfinder` | Broad passive subdomain collection with source attribution | Third-party passive sources | Historical candidates until resolved |
+| urlscan.io | Archived browser requests, screenshots and shared identifiers | Third-party archive; a new scan visits the target | One browser run from one location |
+| Shodan | Favicon, certificate and response pivots across internet scans | Third-party scan index | What Shodan observed at scan time |
+| `httpx` | Live status, title, hashes, JARM and screenshots | Direct requests to every supplied host | Active observation from your address |
+| Wayback CDX | Historical URLs, capture dates and content digests | Internet Archive | Archived captures only; gaps are expected |
+
## Tool reference
| Tool | What it does | Cost |
@@ -467,7 +496,7 @@ question is who runs it and what else they run.
`domain:example-news-daily.com` returns existing public scans, including the third-party requests the page makes. One is a Google
Analytics beacon carrying a property ID.
5. **Pivot on the ID.** That ID searched in PublicWWW returns eleven other domains embedding it,
- and the urlscan query `page.url:* AND "UA-…"` returns nine of the same. Overlapping lists from
+ and the urlscan query `filename:"UA-…"` returns nine of the same. Overlapping lists from
two independent indexes is what makes the network claim defensible.
6. **Confirm the shared template.** `httpx -l hosts.txt -json -favicon -jarm -title` shows the same
favicon hash across the whole set. That hash in `shodan count` returns 14 hosts globally — small
@@ -486,6 +515,13 @@ those people are. Nothing on this page crosses from infrastructure to identity
through [people search](/sheets/osint/people-search) and
[email and phone work](/sheets/osint/email-and-phone).
+What would falsify it: a fresh page capture showing that the analytics or ad identifier was
+injected by a shared third-party template rather than by the operator; certificate history showing
+that the cross-domain certificate was issued to a hosting platform for unrelated customers; or a
+global favicon count large enough to make the hash commonplace. The load-bearing measurements are
+the account-linked analytics and ad identifiers. The shared IP and favicon are corroboration, not
+identity evidence, and must be discarded if either is common outside the eleven-site cluster.
+
## Broader catalogues
- [Domain Name OSINT](https://tools.osintnewsletter.com/tool-categories/domain-name-osint)
diff --git a/src/layouts/Base.astro b/src/layouts/Base.astro
@@ -137,6 +137,30 @@ const canonical = new URL(Astro.url.pathname, Astro.site).href;
<ClientRouter />
</head>
<body>
+ <!-- The signal field: one WebGL canvas fixed to the viewport, under the
+ whole site. First child of <body> so it is the bottom of the stack;
+ the CSS and the layer ladder it belongs to are in chrome.css.
+
+ `transition:persist` is the decision that makes this affordable.
+ The ClientRouter replaces the whole document on every internal
+ navigation, so without it each page would build a WebGL context,
+ compile the shader and throw all of it away a moment later — and
+ the field would restart from a cold phase every time, which reads
+ as a flicker rather than as atmosphere. Persisted, Astro carries
+ this exact element into the new document: one context, one rAF
+ loop, one continuous animation for the whole session.
+
+ Two consequences follow from the element surviving, and both are
+ load-bearing. The `data-bound` guard below is on the canvas
+ itself, so the re-init on `astro:page-load` finds it already
+ marked and does nothing. And there is deliberately NO
+ `astro:before-swap` teardown: the swap is not an unmount here, and
+ tearing down would destroy the context the surviving element still
+ owns. Contrast SectionBanner.astro's fuzz canvas, which does not
+ persist — it is rebuilt per page, so it *must* cancel its rAF and
+ drop its listeners on before-swap or every navigation would leak
+ another loop. -->
+ <canvas class="signal-field" data-signal aria-hidden="true" transition:persist></canvas>
<!-- Shared arrowhead for every inline-SVG flowchart (src/styles/flow.css).
Defined once here so a diagram only has to reference `url(#flow-arrow)`.
`context-stroke` makes the head take each edge's own stroke colour, so
@@ -167,6 +191,42 @@ const canonical = new URL(Astro.url.pathname, Astro.site).href;
<SearchModal />
<script>
import '../scripts/app.ts';
+ import { mountSignalField } from '../scripts/signal-field';
+
+ /**
+ * Bind the signal field. This block imports, so Astro bundles it
+ * rather than treating it as inline source: it ships as an external
+ * module with nothing for the CSP to hash, and if the build does
+ * inline it, scripts/csp-hashes.mjs reads the hash out of the built
+ * HTML. Either way there is no hash to maintain by hand — unlike the
+ * two `is:inline` bootstraps in the head.
+ *
+ * The guard lives on the canvas, which is `transition:persist`ed, so
+ * the element the first load bound is the same element every later
+ * `astro:page-load` sees: it is already marked and the loop keeps
+ * running untouched. The `readyState` pair is the house init shape
+ * (app.ts, SectionBanner.astro) — it covers the first load, where a
+ * deferred module can run either side of `astro:page-load`.
+ *
+ * The teardown the module returns is deliberately dropped. There is
+ * no unmount point for a canvas that outlives every navigation, and
+ * the one case that does need to end — no WebGL, a shader that will
+ * not compile, a lost context — the module handles itself by tearing
+ * down and removing the canvas. The page ground is correct without
+ * it, so there is nothing here to fall back to.
+ */
+ function initSignalField(): void {
+ document
+ .querySelectorAll<HTMLCanvasElement>('[data-signal]:not([data-bound])')
+ .forEach((cv) => {
+ cv.dataset.bound = '1';
+ mountSignalField(cv);
+ });
+ }
+
+ document.addEventListener('astro:page-load', initSignalField);
+ if (document.readyState !== 'loading') initSignalField();
+ else document.addEventListener('DOMContentLoaded', initSignalField);
</script>
</body>
</html>
diff --git a/src/scripts/signal-field.ts b/src/scripts/signal-field.ts
@@ -0,0 +1,750 @@
+/**
+ * The site's atmosphere, drawn in WebGL on one viewport-fixed canvas that
+ * sits behind every page: a network map. Hosts sit on an irregular
+ * lattice, drift a little on their own and are pushed away from the
+ * pointer, so the map bulges around the cursor and settles when it
+ * leaves. The hosts within the cursor's reach lock in the archive's REC
+ * red and the links between them light, so pointing at the page draws an
+ * attack path across it.
+ *
+ * Ported from the main site's `ambient-signal.tsx` (React + WebGL). One
+ * function, one canvas, its own teardown — the caller owns the element
+ * and the lifecycle; this file owns everything that happens on the GPU.
+ *
+ * ── Styles ──
+ * Four tellings of the same field. The site ships the constellation
+ * (`DEFAULT_STYLE`); `?field=N` on the URL swaps in another, for
+ * comparing them on the live site. All four stay in the fragment shader:
+ * the `st()` branches are uniform-constant, so they are coherent across
+ * every fragment and cost effectively nothing, and a byte-identical
+ * shader stays diffable against the original when the main site moves.
+ * 1 mesh hairline lattice with probes running the links
+ * 2 constellation a dense field of plain dots; links only appear near
+ * the cursor, and the cursor itself reaches out to them
+ * 3 circuit links routed at right angles, hosts as pads
+ * 4 radar soft hosts that ping as the cursor passes
+ *
+ * ── The clearing ──
+ * An element carrying `data-index-zone` asks the field to part around it:
+ * every host inside its box is pushed out through the nearest edge to a
+ * margin beyond it, and links, hosts and locks fade to nothing inside.
+ * Nothing on this site carries the attribute yet, so `u_zone` stays at
+ * its empty sentinel and the keep-out is a no-op — see `findZone` below
+ * for why the mechanism is kept anyway.
+ *
+ * ── Colour ──
+ * No palette lives in the shader. The ink and the three accents are read
+ * off the canvas's own computed style, from the theme-relative tokens:
+ * `--text` and the `-soft` washes. Those invert with `[data-theme]` on
+ * `:root` and are re-pinned by `.plate` / `.plate-dawn` for their
+ * subtrees, so the cascade has already answered the question and the
+ * canvas only has to ask the right element — exactly how the banners'
+ * fuzz field resolves its own palette. The `-soft` washes stand in for
+ * the accents because the full-strength inks are set for text and go
+ * muddy as a wash. The canvas is transparent and composites over
+ * whatever ground it is given, so the same program is right on cream and
+ * on ink. Re-read on `daemonmodechange`, because a GPU uniform does not
+ * inherit a CSS variable.
+ *
+ * ── Cost ──
+ * The backing store is capped at device-pixel 1: the lines are a pixel
+ * wide and the fragment cost is per pixel. Layout is never read inside a
+ * frame or a pointer handler — the canvas box and the clearing are
+ * measured by observers into records — so neither a frame nor a pointer
+ * move can force a synchronous layout against whatever else the page is
+ * writing. The pointer is listened for on `window` rather than on the
+ * canvas, so the content in front keeps its hover and the canvas can
+ * stay `pointer-events: none`.
+ *
+ * ── Motion ──
+ * Under `prefers-reduced-motion: reduce` the loop never starts: one frame
+ * is drawn at a fixed phase and redrawn only when the canvas resizes or
+ * the mode flips. The frame also pauses while the tab is hidden. With no
+ * pointer on the page the cursor wanders the field on its own, so a touch
+ * screen still sees hosts being locked.
+ *
+ * ── Degradation ──
+ * No WebGL, a shader that will not compile, a program that will not link
+ * or a lost context all end the same way: the canvas is taken out of the
+ * DOM and the page keeps the ground it already had. The main site falls
+ * back to a second component here because the field is the only thing on
+ * that ground; this page's ground is already correct without it, so
+ * nothing is ported to stand in.
+ */
+
+const VERT = `
+attribute vec2 a_pos;
+void main() { gl_Position = vec4(a_pos, 0.0, 1.0); }
+`;
+
+const FRAG = `
+precision highp float;
+uniform vec2 u_res;
+uniform float u_time;
+uniform vec2 u_mouse;
+uniform float u_scroll;
+uniform float u_style;
+/* The index rows' box in field units (x0, y0, x1, y1), or empty
+ (x1 <= x0) when nothing is asking the field to part. */
+uniform vec4 u_zone;
+uniform vec3 u_ink;
+uniform vec3 u_love;
+uniform vec3 u_foam;
+uniform vec3 u_iris;
+
+float hash(vec2 p) { return fract(sin(dot(p, vec2(127.1, 311.7))) * 43758.5453); }
+float hash3(vec2 p, float k) { return hash(p + vec2(k * 0.731, k * 1.129)); }
+
+bool st(int n) { return int(u_style + 0.5) == n; }
+
+/* Cells per unit of height: ~105px at a 808px hero, or ~80px for the
+ constellation, which is the dense one. */
+float S() { return st(2) ? 10.2 : 7.7; }
+
+/* The cursor, in field units; set once per fragment before the hosts
+ are placed, because a host's position depends on it. */
+vec2 g_cursor = vec2(0.0);
+/* The keep-out box, in field units: the rows' box shifted with the
+ field's scroll offset. Set beside the cursor. */
+vec4 g_zone = vec4(0.0, 0.0, -1.0, -1.0);
+
+/* Signed distance to the keep-out box: negative inside. */
+float zoneSd(vec2 p) {
+ if (g_zone.z <= g_zone.x) return 1e3;
+ vec2 c = 0.5 * (g_zone.xy + g_zone.zw);
+ vec2 h = 0.5 * (g_zone.zw - g_zone.xy);
+ vec2 q = abs(p - c) - h;
+ return length(max(q, 0.0)) + min(max(q.x, q.y), 0.0);
+}
+
+/* The host in cell c, or far off-field if the cell is empty. Each host
+ drifts a little on its own, and the cursor pushes the ones near it
+ away with a soft falloff, so the links stretch and the map bulges
+ around the pointer and settles again when it leaves. */
+vec2 host(vec2 c) {
+ float keep = st(2) ? 0.9 : (st(4) ? 0.70 : 0.78);
+ if (hash(c) > keep) return vec2(-1e3);
+ vec2 h = (c + 0.22 + 0.56 * vec2(hash3(c, 1.0), hash3(c, 2.0))) / S();
+ float k1 = hash3(c, 3.0);
+ float k2 = hash3(c, 4.0);
+ float drift = st(2) ? 0.012 : 0.011;
+ h += drift * vec2(sin(u_time * (0.28 + k1 * 0.4) + k1 * 6.2832),
+ cos(u_time * (0.22 + k2 * 0.4) + k2 * 6.2832));
+ vec2 d = h - g_cursor;
+ float dist = length(d);
+ float push = smoothstep(0.36, 0.0, dist) * 0.07;
+ h += (d / max(dist, 1e-4)) * push;
+ /* out of the keep-out box, through its nearest edge, to a margin
+ beyond it — so the rows have clear ground and a clear surround */
+ float margin = 0.04;
+ float sd = zoneSd(h);
+ if (sd < margin) {
+ vec2 c2 = 0.5 * (g_zone.xy + g_zone.zw);
+ vec2 hh = 0.5 * (g_zone.zw - g_zone.xy);
+ vec2 q = abs(h - c2) - hh;
+ vec2 dir = (q.x > q.y) ? vec2(sign(h.x - c2.x), 0.0) : vec2(0.0, sign(h.y - c2.y));
+ h += dir * (margin - sd);
+ }
+ return h;
+}
+
+/* Distance from p to segment ab, and the position along it. */
+vec2 seg(vec2 p, vec2 a, vec2 b) {
+ vec2 ab = b - a;
+ float t = clamp(dot(p - a, ab) / max(1e-6, dot(ab, ab)), 0.0, 1.0);
+ return vec2(length(p - (a + ab * t)), t);
+}
+
+/* Whether a link exists between cell a and its neighbour in direction d. */
+bool linked(vec2 a, vec2 d) {
+ float h = hash(a * 3.1 + d * 17.7);
+ if (st(3)) return (abs(d.x) + abs(d.y) > 1.5) ? h > 0.85 : h > 0.38;
+ return (abs(d.x) + abs(d.y) > 1.5) ? h > 0.72 : h > 0.42;
+}
+
+/* The output, built up as premultiplied colour over a clear ground:
+ each layer is composited "over" what is there, which on an opaque
+ ground is exactly the mix() chain it replaces. */
+vec4 g_acc = vec4(0.0);
+void lay(vec3 c, float a) {
+ a = clamp(a, 0.0, 1.0);
+ g_acc = vec4(g_acc.rgb * (1.0 - a) + c * a, g_acc.a * (1.0 - a) + a);
+}
+
+void main() {
+ float aspect = u_res.x / u_res.y;
+ vec2 uv = (gl_FragCoord.xy - 0.5 * u_res) / u_res.y;
+ vec2 mu = u_mouse * vec2(aspect * 0.5, 0.5);
+ /* The field slides up a little as the page is scrolled past. */
+ vec2 fp = uv + vec2(0.0, u_scroll * 0.12);
+ vec2 fm = mu + vec2(0.0, u_scroll * 0.12);
+ g_cursor = fm;
+ g_zone = u_zone + vec4(0.0, u_scroll * 0.12, 0.0, u_scroll * 0.12);
+
+ float px = 1.0 / u_res.y;
+
+ /* The 64px hairline grid the signal map draws; the circuit takes a
+ finer 32px one. The constellation and the radar have none. */
+ if (st(1) || st(3)) {
+ float cellPx = st(3) ? 32.0 : 64.0;
+ vec2 gp = gl_FragCoord.xy / cellPx;
+ vec2 gq = abs(fract(gp) - 0.5);
+ float grid = 1.0 - smoothstep(0.0, 1.5 / cellPx, min(0.5 - gq.x, 0.5 - gq.y));
+ lay(u_ink, grid * (st(3) ? 0.035 : 0.045));
+ }
+
+ float reach = st(2) ? 0.28 : 0.24;
+ vec2 cell = floor(fp * S());
+ float linkA = 0.0; /* resting link ink */
+ float pathA = 0.0; /* links lit by the lock */
+ float probeA = 0.0; /* probes on the links */
+ float hostA = 0.0; /* resting host dots */
+ float haloA = 0.0; /* the soft glow around a host */
+ float lockA = 0.0; /* locked hosts */
+ float ringA = 0.0; /* the lock ring */
+ float grabA = 0.0; /* constellation: cursor-to-host lines */
+ float pingA = 0.0; /* radar: expanding rings */
+
+ /* Two cells each way, because a pushed host can leave its own cell;
+ anything farther than that cannot reach this pixel and is skipped
+ before its links are measured. */
+ for (int j = -2; j <= 2; j++) {
+ for (int i = -2; i <= 2; i++) {
+ vec2 a = cell + vec2(float(i), float(j));
+ vec2 ha = host(a);
+ if (ha.x < -100.0) continue;
+ if (length(ha - fp) > 2.3 / S()) continue;
+
+ float toCursor = length(ha - fm);
+ float la = smoothstep(reach, reach * 0.55, toCursor);
+
+ /* The constellation only shows links near the cursor, so links
+ whose near end is well out of reach are not measured at all —
+ which is most of them, on most of the screen. */
+ bool links = !(st(2) && toCursor > reach + 2.0 / S());
+
+ if (links) for (int k = 0; k < 4; k++) {
+ vec2 d = (k == 0) ? vec2(1.0, 0.0) : (k == 1) ? vec2(0.0, 1.0)
+ : (k == 2) ? vec2(1.0, 1.0) : vec2(1.0, -1.0);
+ if (!linked(a, d)) continue;
+ vec2 hb = host(a + d);
+ if (hb.x < -100.0) continue;
+ float lb = smoothstep(reach, reach * 0.55, length(hb - fm));
+ float lit = la * lb;
+
+ float along = 0.0;
+ float dist = 0.0;
+ if (st(3)) {
+ /* routed at a right angle through a corner, the corner's side
+ chosen per link so the routes do not all turn the same way */
+ vec2 corner = hash(a * 7.7 + d) > 0.5 ? vec2(hb.x, ha.y) : vec2(ha.x, hb.y);
+ vec2 s1 = seg(fp, ha, corner);
+ vec2 s2 = seg(fp, corner, hb);
+ float l1 = length(corner - ha);
+ float l2 = length(hb - corner);
+ float L = max(1e-5, l1 + l2);
+ if (s1.x < s2.x) { dist = s1.x; along = s1.y * l1 / L; }
+ else { dist = s2.x; along = (l1 + s2.y * l2) / L; }
+ } else {
+ vec2 s = seg(fp, ha, hb);
+ dist = s.x;
+ along = s.y;
+ }
+ float line = 1.0 - smoothstep(0.0, 1.4 * px, dist - 0.35 * px);
+ linkA = max(linkA, line);
+ pathA = max(pathA, line * lit);
+
+ /* a probe runs the link if the link carries traffic */
+ float eh = hash(a * 5.3 + d * 2.9);
+ if ((st(1) || st(3)) && eh > 0.55) {
+ float dir = eh > 0.78 ? 1.0 : -1.0;
+ float pos = fract(u_time * (0.10 + eh * 0.08) * dir + eh * 7.0);
+ float dp = along - pos;
+ float head = exp(-dp * dp * 2200.0);
+ float tail = (dir > 0.0)
+ ? ((dp < 0.0) ? exp(dp * 14.0) : 0.0)
+ : ((dp > 0.0) ? exp(-dp * 14.0) : 0.0);
+ float on = 1.0 - smoothstep(0.0, 2.2 * px, dist);
+ probeA = max(probeA, on * (head + tail * 0.35) * (0.35 + 0.65 * lit));
+ }
+ }
+
+ /* the host itself */
+ vec2 rel = fp - ha;
+ float dh = st(3) ? max(abs(rel.x), abs(rel.y)) : length(rel);
+ float r0 = st(4) ? 2.0 * px : 1.2 * px;
+ hostA = max(hostA, 1.0 - smoothstep(r0, r0 + 1.2 * px, dh));
+ if (st(4)) haloA = max(haloA, exp(-dh * dh * 9000.0));
+ lockA = max(lockA, (1.0 - smoothstep(r0 + 0.4 * px, r0 + 1.8 * px, dh)) * la);
+ float ring = abs(dh - 7.0 * px);
+ ringA = max(ringA, (1.0 - smoothstep(0.0, 1.3 * px, ring)) * la);
+
+ if (st(2) && la > 0.0) {
+ /* the cursor reaches out to what it can see */
+ vec2 g = seg(fp, fm, ha);
+ float gl = 1.0 - smoothstep(0.0, 1.4 * px, g.x - 0.35 * px);
+ grabA = max(grabA, gl * la * (0.4 + 0.6 * (1.0 - g.y)));
+ }
+ if (st(4)) {
+ /* each host pings while the cursor is near, and now and then on
+ its own; the ring grows out and fades */
+ float ph = hash3(a, 9.0);
+ float idle = step(0.86, ph) * 0.35;
+ float amp = max(la, idle);
+ float cyc = fract(u_time * (0.35 + ph * 0.25) + ph * 3.0);
+ float rr = cyc * 0.11;
+ float band = abs(length(rel) - rr);
+ float pr = (1.0 - smoothstep(0.0, 1.2 * px, band)) * (1.0 - cyc) * amp;
+ pingA = max(pingA, pr);
+ }
+ }
+ }
+
+ /* Nothing of the map is drawn inside the keep-out box; it fades
+ over the last few pixels before the edge. */
+ float keep = smoothstep(0.0, 0.02, zoneSd(fp));
+ linkA *= keep; pathA *= keep; probeA *= keep; hostA *= keep;
+ haloA *= keep; lockA *= keep; ringA *= keep; grabA *= keep; pingA *= keep;
+
+ if (st(1)) {
+ lay(u_ink, linkA * 0.075);
+ lay(u_love, pathA * 0.6);
+ lay(u_foam, probeA * 0.85);
+ lay(u_ink, hostA * 0.32);
+ } else if (st(2)) {
+ lay(u_ink, linkA * 0.03);
+ lay(u_love, pathA * 0.5);
+ lay(u_iris, grabA * 0.45);
+ lay(u_ink, hostA * 0.38);
+ } else if (st(3)) {
+ lay(u_ink, linkA * 0.09);
+ lay(u_love, pathA * 0.6);
+ lay(u_foam, probeA * 0.9);
+ lay(u_ink, hostA * 0.4);
+ } else {
+ lay(u_ink, linkA * 0.05);
+ lay(u_love, pathA * 0.45);
+ lay(u_foam, haloA * 0.22);
+ lay(u_foam, pingA * 0.6);
+ lay(u_ink, hostA * 0.45);
+ }
+ lay(u_love, lockA * 0.95);
+ lay(u_love, ringA * 0.7);
+
+ /* The pointer's own wash: a wide, faint iris. */
+ vec2 mp = uv - mu;
+ lay(u_iris, exp(-dot(mp, mp) * 9.0) * 0.07);
+
+ gl_FragColor = g_acc;
+}
+`;
+
+/** One of the four tellings; `?field=N` picks between them. */
+type SignalStyle = 1 | 2 | 3 | 4;
+
+/** The style the site ships with: the constellation. */
+const DEFAULT_STYLE: SignalStyle = 2;
+
+const TOKENS = ['ink', 'love', 'foam', 'iris'] as const;
+type Token = (typeof TOKENS)[number];
+type Rgb = [number, number, number];
+
+/** #rrggbb to 0–1 channels. Anything else gives the fallback. */
+function hexToRgb(hex: string, fallback: Rgb): Rgb {
+ const m = /^#?([0-9a-f]{6})$/i.exec(hex.trim());
+ if (!m) return fallback;
+ const n = parseInt(m[1], 16);
+ return [((n >> 16) & 255) / 255, ((n >> 8) & 255) / 255, (n & 255) / 255];
+}
+
+/** The night palette, for when a token is missing or is not a plain hex. */
+const FALLBACK: Record<Token, Rgb> = {
+ ink: hexToRgb('#e0def4', [0.88, 0.87, 0.96]),
+ love: hexToRgb('#eb6f92', [0.92, 0.44, 0.57]),
+ foam: hexToRgb('#9ccfd8', [0.61, 0.81, 0.85]),
+ iris: hexToRgb('#c4a7e7', [0.77, 0.65, 0.9]),
+};
+
+/**
+ * The custom property each uniform reads. Theme-relative on purpose, not
+ * the absolute `--color-night-*` / `--color-dawn-*` set: these four
+ * invert with `[data-theme]` on `:root` and are re-pinned by `.plate` and
+ * `.plate-dawn` for their subtrees, so the field follows a retuned plate
+ * without a code change. All four resolve to a plain `#rrggbb`, which is
+ * what `hexToRgb` wants — no `color-mix()` values are involved.
+ */
+const SOURCES: Record<Token, string> = {
+ ink: '--text',
+ love: '--love-soft',
+ foam: '--foam-soft',
+ iris: '--iris-soft',
+};
+
+function readPalette(el: HTMLElement): Record<Token, Rgb> {
+ const cs = getComputedStyle(el);
+ const out = {} as Record<Token, Rgb>;
+ for (const t of TOKENS) out[t] = hexToRgb(cs.getPropertyValue(SOURCES[t]), FALLBACK[t]);
+ return out;
+}
+
+function compile(gl: WebGLRenderingContext, type: number, src: string): WebGLShader | null {
+ const sh = gl.createShader(type);
+ if (!sh) return null;
+ gl.shaderSource(sh, src);
+ gl.compileShader(sh);
+ if (!gl.getShaderParameter(sh, gl.COMPILE_STATUS)) {
+ gl.deleteShader(sh);
+ return null;
+ }
+ return sh;
+}
+
+/**
+ * `?field=N` on the URL overrides the shipped style, for comparing the
+ * four side by side on the live site. Anything else is the default.
+ */
+function styleFromUrl(): SignalStyle | null {
+ try {
+ const n = Number(new URLSearchParams(window.location.search).get('field'));
+ return n === 1 || n === 2 || n === 3 || n === 4 ? n : null;
+ } catch {
+ return null;
+ }
+}
+
+/** The phase the still frame is drawn at: probes spread, hosts settled. */
+const HELD_TIME = 31.0;
+/** How long the pointer may rest before the cursor starts to wander. */
+const IDLE_MS = 2800;
+
+/**
+ * Bind the field to `canvas` and start it. Returns the teardown: it
+ * cancels the frame, drops every listener and observer, and deletes the
+ * GL objects. Safe to call twice — the context-lost handler tears down on
+ * its own, and the caller's `astro:before-swap` will still fire.
+ *
+ * If WebGL is unavailable, or the program will not build, the canvas is
+ * removed from the DOM and the teardown is a no-op. That is the whole
+ * degradation signal: the element the caller handed over is gone, and
+ * nothing is left running.
+ */
+export function mountSignalField(canvas: HTMLCanvasElement): () => void {
+ const noop = (): void => {};
+
+ /* Nothing to fall back to and nothing to say: the page's ground is
+ already right without the field, so the canvas simply leaves. */
+ const degrade = (): void => { canvas.remove(); };
+
+ const attrs: WebGLContextAttributes = {
+ alpha: true,
+ premultipliedAlpha: true,
+ antialias: false,
+ depth: false,
+ stencil: false,
+ powerPreference: 'low-power',
+ };
+ const gl =
+ canvas.getContext('webgl', attrs) ||
+ (canvas.getContext('experimental-webgl', attrs) as WebGLRenderingContext | null);
+ if (!gl) {
+ degrade();
+ return noop;
+ }
+
+ const vs = compile(gl, gl.VERTEX_SHADER, VERT);
+ const fs = compile(gl, gl.FRAGMENT_SHADER, FRAG);
+ const prog = gl.createProgram();
+ /* Whatever did survive is released here rather than left for the
+ context to collect: React unmounted the whole canvas on this path, so
+ the original never had to, but here the caller keeps its page. */
+ if (!vs || !fs || !prog) {
+ if (vs) gl.deleteShader(vs);
+ if (fs) gl.deleteShader(fs);
+ if (prog) gl.deleteProgram(prog);
+ degrade();
+ return noop;
+ }
+ gl.attachShader(prog, vs);
+ gl.attachShader(prog, fs);
+ gl.linkProgram(prog);
+ if (!gl.getProgramParameter(prog, gl.LINK_STATUS)) {
+ gl.deleteProgram(prog);
+ gl.deleteShader(vs);
+ gl.deleteShader(fs);
+ degrade();
+ return noop;
+ }
+ gl.useProgram(prog);
+
+ const buf = gl.createBuffer();
+ gl.bindBuffer(gl.ARRAY_BUFFER, buf);
+ gl.bufferData(
+ gl.ARRAY_BUFFER,
+ new Float32Array([-1, -1, 1, -1, -1, 1, 1, 1]),
+ gl.STATIC_DRAW,
+ );
+ const aPos = gl.getAttribLocation(prog, 'a_pos');
+ gl.enableVertexAttribArray(aPos);
+ gl.vertexAttribPointer(aPos, 2, gl.FLOAT, false, 0, 0);
+ gl.clearColor(0, 0, 0, 0);
+
+ const u = (name: string): WebGLUniformLocation | null => gl.getUniformLocation(prog, name);
+ const uRes = u('u_res');
+ const uTime = u('u_time');
+ const uMouse = u('u_mouse');
+ const uScroll = u('u_scroll');
+ const uZone = u('u_zone');
+ const uCol = {} as Record<Token, WebGLUniformLocation | null>;
+ for (const t of TOKENS) uCol[t] = u(`u_${t}`);
+ gl.uniform1f(u('u_style'), styleFromUrl() ?? DEFAULT_STYLE);
+ gl.uniform4f(uZone, 0, 0, -1, -1);
+
+ let palette = readPalette(canvas);
+ const pushPalette = (): void => {
+ for (const t of TOKENS) gl.uniform3fv(uCol[t], palette[t]);
+ };
+ pushPalette();
+
+ /* ── State the frame reads ── */
+ const mouse = { x: 0.3, y: 0.1 }; // where the cursor should be, −1…1
+ const eased = { x: 0.3, y: 0.1 }; // where the shader thinks it is
+ let lastMove = -Infinity;
+ let scroll = 0;
+ let raf = 0;
+ let running = false;
+ let visible = !document.hidden;
+ let lost = false;
+ const live = !matchMedia('(prefers-reduced-motion: reduce)').matches;
+ const start = performance.now();
+
+ /* The canvas box, measured by the observers below and never inside a
+ frame or a pointer event: a layout read in either place forces the
+ browser to flush pending style and layout synchronously.
+
+ The original stored the box's *page*-space top, because its canvas
+ scrolled with the page: the viewport-relative top changed on every
+ scroll while the page-relative one did not. This canvas is fixed to
+ the viewport, so that inverts — the viewport-relative top is the
+ constant, and the scroll offset must not be added back in. Pointer
+ coordinates are viewport-relative already, so the conversion below
+ gets simpler rather than harder. */
+ const box = { left: 0, top: 0, width: 1, height: 1 };
+ const measureBox = (): void => {
+ const r = canvas.getBoundingClientRect();
+ box.left = r.left;
+ box.top = r.top;
+ box.width = r.width || 1;
+ box.height = r.height || 1;
+ };
+
+ const size = (): void => {
+ /* Device-pixel 1, whatever the display: the lines are a pixel wide,
+ the field is a background, and every extra pixel is a host loop. */
+ const w = Math.max(1, Math.round(canvas.clientWidth));
+ const h = Math.max(1, Math.round(canvas.clientHeight));
+ if (canvas.width !== w || canvas.height !== h) {
+ canvas.width = w;
+ canvas.height = h;
+ }
+ gl.viewport(0, 0, w, h);
+ gl.uniform2f(uRes, w, h);
+ };
+
+ /* The clearing: an element with `data-index-zone`, measured against the
+ canvas whenever either changes shape. Nothing on this site carries
+ the attribute, so this resolves to nothing and `u_zone` keeps its
+ empty sentinel — `zoneSd` then returns on its first line and the
+ keep-out costs one uniform per fragment.
+
+ It is kept rather than stripped because it is load-bearing behaviour
+ in the original, and because the home page's index section is the
+ structural analogue of the rows it was written for: if the field is
+ ever asked to part around them, this is what it will need. The one
+ thing to redo first is the measurement — a viewport-fixed canvas and
+ a zone that scrolls with the page part company on every scroll, so
+ the box would have to be re-measured against the scroll offset
+ rather than only when either element changes shape.
+
+ The lookup is retried a few times rather than once: the original's
+ hero mounted its rows after it had measured itself, and anything
+ here that marks a zone from script would do the same. */
+ let zoneEl: HTMLElement | null = null;
+ let zoneRo: ResizeObserver | null = null;
+ const measureZone = (): void => {
+ if (!zoneEl || !zoneEl.isConnected) {
+ gl.uniform4f(uZone, 0, 0, -1, -1);
+ return;
+ }
+ const c = canvas.getBoundingClientRect();
+ const z = zoneEl.getBoundingClientRect();
+ const H = c.height || 1;
+ gl.uniform4f(
+ uZone,
+ (z.left - c.left - c.width / 2) / H,
+ (c.height / 2 - (z.bottom - c.top)) / H,
+ (z.right - c.left - c.width / 2) / H,
+ (c.height / 2 - (z.top - c.top)) / H,
+ );
+ };
+ let zoneTimer = 0;
+ const findZone = (attempt = 0): void => {
+ const el = document.querySelector<HTMLElement>('[data-index-zone]');
+ if (el) {
+ zoneEl = el;
+ zoneRo = new ResizeObserver(measureZone);
+ zoneRo.observe(el);
+ measureZone();
+ return;
+ }
+ if (attempt < 8) zoneTimer = window.setTimeout(() => findZone(attempt + 1), 250);
+ };
+
+ const draw = (t: number): void => {
+ if (lost) return;
+ gl.uniform1f(uTime, t);
+ gl.uniform2f(uMouse, eased.x, eased.y);
+ gl.uniform1f(uScroll, scroll);
+ gl.clear(gl.COLOR_BUFFER_BIT);
+ gl.drawArrays(gl.TRIANGLE_STRIP, 0, 4);
+ };
+
+ const still = (): void => {
+ size();
+ eased.x = 0.3;
+ eased.y = 0.1;
+ draw(HELD_TIME);
+ };
+
+ const tick = (now: number): void => {
+ raf = requestAnimationFrame(tick);
+ const t = (now - start) / 1000;
+ /* Left alone, the cursor walks a slow figure across the field — a
+ touch screen never sends a pointer, and the lock is the piece's
+ whole gesture. */
+ if (now - lastMove > IDLE_MS) {
+ mouse.x = 0.45 * Math.sin(t * 0.19) + 0.15 * Math.sin(t * 0.47);
+ mouse.y = 0.4 * Math.cos(t * 0.14) + 0.12 * Math.cos(t * 0.37);
+ }
+ /* A 0.07 step settles in about a third of a second: weight, not lag. */
+ eased.x += (mouse.x - eased.x) * 0.07;
+ eased.y += (mouse.y - eased.y) * 0.07;
+ draw(t);
+ };
+
+ /* One gate for the loop, as in the original — minus its in-view test.
+ The original observed its own canvas with an IntersectionObserver and
+ paused when the hero scrolled away; a canvas fixed to the viewport is
+ always intersecting, so that observer could only ever report true.
+ It is dropped deliberately rather than quietly: the tab-hidden pause
+ below is the one that still does work here. */
+ const sync = (): void => {
+ const want = live && visible;
+ if (want && !running) {
+ running = true;
+ size();
+ raf = requestAnimationFrame(tick);
+ } else if (!want && running) {
+ running = false;
+ cancelAnimationFrame(raf);
+ }
+ if (!live) still();
+ };
+
+ /* The pointer, relative to the canvas, −1…1 on both axes, placed from
+ the measured box rather than a fresh rect. Listened for on the window
+ because the canvas does not take events — the content in front of it
+ must keep its hover. */
+ const onMove = (e: PointerEvent): void => {
+ if (e.pointerType === 'touch') return;
+ const x = ((e.clientX - box.left) / box.width) * 2 - 1;
+ const y = -(((e.clientY - box.top) / box.height) * 2 - 1);
+ /* Off the field the cursor is handed back to the wanderer. */
+ if (x < -1.05 || x > 1.05 || y < -1.05 || y > 1.05) return;
+ mouse.x = x;
+ mouse.y = y;
+ lastMove = e.timeStamp;
+ };
+ /* How far the page has been scrolled, in viewport heights — the field
+ slides up a little against it and then holds. `box.height` is the
+ viewport now rather than a hero's band, so one screen of scrolling
+ spends the whole term. */
+ const onScroll = (): void => {
+ scroll = Math.max(0, Math.min(1, window.scrollY / box.height));
+ };
+ const onVis = (): void => {
+ visible = !document.hidden;
+ sync();
+ };
+ /* A GPU uniform does not inherit. The variables have already changed by
+ the time this fires, so re-reading them is all it takes — but a still
+ frame has to be redrawn by hand, or the field keeps the old palette
+ until something else resizes it. */
+ const onMode = (): void => {
+ palette = readPalette(canvas);
+ pushPalette();
+ if (!running) still();
+ };
+ const onLost = (e: Event): void => {
+ e.preventDefault();
+ lost = true;
+ teardown();
+ degrade();
+ };
+
+ const ro = new ResizeObserver(() => {
+ measureBox();
+ size();
+ measureZone();
+ if (!running) still();
+ });
+ ro.observe(canvas);
+
+ /* No pointer listener under reduced motion: there is no loop to read
+ the cursor, and the still frame is drawn at a fixed one. */
+ if (live) window.addEventListener('pointermove', onMove, { passive: true });
+ window.addEventListener('scroll', onScroll, { passive: true });
+ window.addEventListener('resize', measureBox);
+ document.addEventListener('visibilitychange', onVis);
+ window.addEventListener('daemonmodechange', onMode);
+ canvas.addEventListener('webglcontextlost', onLost);
+
+ /* An arrow rather than a declaration, and placed here rather than
+ beside the handler that calls it: a hoisted `function` could in
+ principle run before `gl` was narrowed, so the compiler widens it
+ back to possibly-null inside the body. `onLost` reaches forward to
+ this binding, which is sound because it can only be called once the
+ listener above it has been bound. */
+ let torn = false;
+ const teardown = (): void => {
+ if (torn) return;
+ torn = true;
+ running = false;
+ cancelAnimationFrame(raf);
+ clearTimeout(zoneTimer);
+ ro.disconnect();
+ zoneRo?.disconnect();
+ window.removeEventListener('pointermove', onMove);
+ window.removeEventListener('scroll', onScroll);
+ window.removeEventListener('resize', measureBox);
+ document.removeEventListener('visibilitychange', onVis);
+ window.removeEventListener('daemonmodechange', onMode);
+ canvas.removeEventListener('webglcontextlost', onLost);
+ gl.deleteBuffer(buf);
+ gl.deleteProgram(prog);
+ gl.deleteShader(vs);
+ gl.deleteShader(fs);
+ /* The context is left to the GC rather than lost by hand: a forced
+ loss fires `webglcontextlost` after this cleanup has run, and the
+ next mount of the same canvas would hear it as a real loss and
+ degrade for no reason. That matters more here than in the original,
+ because a view-transition navigation unmounts and mounts in quick
+ succession. */
+ };
+
+ measureBox();
+ onScroll();
+ findZone();
+ sync();
+
+ return teardown;
+}
diff --git a/src/styles/chrome.css b/src/styles/chrome.css
@@ -9,6 +9,7 @@
.slab flat panel — 1px rule, square, no shadow
.plate inverse (night) panel, defined in tokens.css
.grain the film over everything
+ .signal-field the dot field under everything
.micro 10px mono label — the site's furniture type
.eyebrow `03 ^: LABEL` lockup with a keyed rule
.btn flat control
@@ -47,6 +48,49 @@
z-index: 2;
}
+/* ---- The field ----------------------------------------------------------- */
+/* The other sheet pinned to the viewport, and the film's opposite number:
+ one WebGL canvas under the whole site, drawing the signal map (hosts on
+ an irregular lattice, links lighting where the pointer reaches). It is
+ mounted once in Base.astro and lives for the whole session;
+ `src/scripts/signal-field.ts` owns everything else about it.
+
+ It is a ground, not a control, so it takes no events at all — the
+ pointer it follows is listened for on `window`, which leaves every link
+ and button on the page its own hover.
+
+ The fixed layers, bottom to top, are the site's whole stacking contract
+ and the reason this one sits at 0:
+
+ 0 .signal-field this canvas
+ 1 .grain the film — the canvas is part of its backdrop,
+ so the field is grained along with the page
+ 2 main, .site-header, .site-footer all page content (global.css)
+ 40 .to-top
+ 60 .site-header once stuck, and `.mast` (daemon.css)
+ 90 .scroll-progress
+ 100 .search-overlay
+
+ 0 rather than -1 because the ladder already leaves 0 free, and because
+ "the bottom layer of the document" is what this is — not something
+ hiding behind it. Either index is in fact visible, and the reason is
+ worth writing down, since every page paints an opaque ground: `html`
+ sets no background, so `body`'s `--bg` propagates to the root element
+ and is painted as the canvas background, *below every child of body*.
+ A fixed child of body therefore paints over the page ground rather
+ than under it. What does hide the field is an opaque box in the
+ content layer — the banner plate, the marquee band, the footer — and
+ that is intended; the field reads in the paper between them. */
+.signal-field {
+ position: fixed;
+ inset: 0;
+ z-index: 0;
+ width: 100%;
+ height: 100%;
+ display: block;
+ pointer-events: none;
+}
+
/* ---- Panels -------------------------------------------------------------- */
.slab {
border: 1px solid var(--rule);