daemon-sec-cheatsheet

The cheatsheet vault for operators: AD, enumeration, exploitation, priv-esc, web, DFIR
git clone https://git.daemon-sec.xyz/daemon-sec-cheatsheet.git
Log | Files | Refs | README | LICENSE

commit d69495a0717a9abaec8ab18ea0cfe3f50cc80cbd
parent 6ea21eb928bb35cb50abfbe09b3d905939de5340
Author: DAEMON <zer0sec.xp@icloud.com>
Date:   Sat,  3 Oct 2026 18:51:29 +0100

Give the OSINT sheets working commands instead of tool catalogues

The 16 OSINT sheets carried 709 tool links and 33 code fences between
them, and only two of them (usernames-and-accounts, email-and-phone) had
per-tool walkthroughs. people-search named 97 tools and contained no
command at all. These eight now follow the shape of the two that were
already right: a numbered Method, per-tool sections with install plus
six to ten commented invocations, a closing note on false positives and
limits, and a worked example that starts from one concrete datum.

Catalogue tables and the Bellingcat / tools.osintnewsletter.com
references are kept intact — this adds depth rather than replacing the
orientation the link lists give.

Flags and endpoints were checked against installed binaries or probed
live rather than recalled, which turned up breakage in the existing
sheets:

- ffprobe -show_entries frame=pkt_pts_time was removed in ffmpeg 5 and
  silently returns an empty column; replaced with pts_time, and
  -vsync vfr with -fps_mode vfr.
- ACLED's api.acleddata.com has no DNS record at all and the key=/email=
  scheme is retired; the sheet's previous curl example was dead.
- ADS-B Exchange retired its freemium API, so adsb.lol is documented as
  the keyless replacement.
- Bluesky post search needs a session (403 keyless) while profile, feed
  and graph reads stay open; amass -passive is a deprecated no-op in v5;
  Overpass 406s without a User-Agent.

Tools that are now paywalled, key-gated or abandoned say so in one line
and name the live alternative, because a cheatsheet that confidently
prints a dead command costs more time than an empty section.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Diffstat:
Msrc/content/sheets/osint/companies-and-finance.md | 530++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-----
Msrc/content/sheets/osint/conflict-and-environment.md | 542++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-----
Msrc/content/sheets/osint/image-video-forensics.md | 396+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++----------
Msrc/content/sheets/osint/maps-and-satellite-imagery.md | 353++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-----
Msrc/content/sheets/osint/people-search.md | 359++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-
Msrc/content/sheets/osint/social-media-platforms.md | 544++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-------
Msrc/content/sheets/osint/transport-tracking.md | 460+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++------
Msrc/content/sheets/osint/websites-and-infrastructure.md | 424++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++---------
8 files changed, 3354 insertions(+), 254 deletions(-)

diff --git a/src/content/sheets/osint/companies-and-finance.md b/src/content/sheets/osint/companies-and-finance.md @@ -4,9 +4,9 @@ description: "Corporate registries, filings and beneficial-ownership data for wo category: osint subcategory: "Corporate & Financial" tags: [osint, companies, finance, ownership, filings] -tools: [opencorporates, edgar, aleph, companies-house] +tools: [opencorporates, edgar, edgartools, aleph, companies-house, opensanctions, icij-offshore-leaks] difficulty: intermediate -updated: 2026-09-28 +updated: 2026-10-03 references: - name: "Bellingcat's Online Investigation Toolkit" url: "https://bellingcat.gitbook.io/toolkit" @@ -26,62 +26,447 @@ references: Establishing what a company is, who owns it, and what it has told regulators. The pattern that matters: the registry tells you the *legal* structure, and the legal structure is often designed to -obscure who benefits. Getting to the human at the end takes several sources. +obscure who benefits. Getting to the human at the end takes several sources in a particular order. ## Method -1. **Start at the national registry**, which is authoritative. Aggregators are convenient but stale. -2. **Get the registration number** and use it thereafter. Company names are reused and change. +1. **Start at the national registry**, which is authoritative. Aggregators are convenient but + stale, and for anything you will publish the registry page is the citation. +2. **Get the registration number** and use it thereafter. Company names are reused across + jurisdictions, reused within them after dissolution, and changed deliberately. 3. **Pull the filings, not the summary.** Annual accounts, director appointments and charges carry - addresses, signatures, auditors and related parties that the summary page omits. + addresses, signatures, auditors and related parties that the summary page omits. The scanned PDF + often has a signature and a handwritten date the structured data does not. 4. **Follow officers sideways.** One director's other appointments frequently reveal the group the - company actually belongs to. -5. **Cross-border means repeat the whole process** in each jurisdiction. A chain terminating in a - secrecy jurisdiction is the normal outcome, not a failure. -6. **Check leak archives** where the chain goes dark — OCCRP Aleph and ICIJ Offshore Leaks hold - what registries do not. + company actually belongs to — and distinguish a decision-maker from a nominee, because a + nominee has dozens or hundreds. +5. **Screen the names you have** against sanctions and politically-exposed-person data before you + go further. A hit changes what the rest of the chain means, and it is cheap to check. +6. **Cross-border means repeating the whole process** in each jurisdiction, and each one publishes + something different: the UK publishes officers and PSCs free, Delaware publishes almost nothing, + the EU publishes through BRIS with per-country gaps, and secrecy jurisdictions publish the + existence of the company and nothing else. A chain terminating in one of them is the normal + outcome, not a failure. +7. **Check leak archives where the chain goes dark.** OCCRP Aleph and ICIJ Offshore Leaks hold what + registries do not, with the caveat that leaked data is a snapshot of one moment. +8. **Write down what each jurisdiction refused to tell you.** "Beneficial owner not published in + this jurisdiction" is a finding about the structure, and it is the sentence that makes the + analysis defensible rather than incomplete. ## Core sources | Source | Coverage | | --- | --- | | [OpenCorporates](https://opencorporates.com/) | The largest open company database, ~200 jurisdictions. Best starting point for "does this company exist and where". | -| [SEC EDGAR](https://www.sec.gov/edgar/search/) | All US public-company filings. Full-text searchable and genuinely free. | -| [EDGAR Suite](https://edgar-suite.vercel.app/) | Friendlier interface over EDGAR for exploring filings. | -| [OCCRP Aleph](https://aleph.occrp.org/) | Registries, leaks and court records in one index. The tool for cross-border work. | -| [ICIJ Offshore Leaks](https://offshoreleaks.icij.org/) | Panama / Paradise / Pandora Papers entities and officers. | +| [Companies House](https://find-and-update.company-information.service.gov.uk/) | UK and Gibraltar. Free, complete filing history including scanned documents, and a clean API. The best-documented major registry. | +| [SEC EDGAR](https://www.sec.gov/edgar/search/) | All US public-company filings. Full-text searchable back to 2001 and genuinely free. | +| [OCCRP Aleph](https://aleph.occrp.org/) | Registries, leaks and court records in one index, with a shared entity model. The tool for cross-border work. | +| [OpenSanctions](https://www.opensanctions.org/) | Consolidated sanctions lists, PEP data and persons of interest, as bulk data and an API. | +| [ICIJ Offshore Leaks](https://offshoreleaks.icij.org/) | Panama / Paradise / Pandora Papers entities, officers and intermediaries. | | [Open Ownership](https://register.openownership.org/) | Beneficial-ownership data, where it is published at all. | +| [BRIS](https://e-justice.europa.eu/content_find_a_company-489-en.do) | Consolidated search across most EU, Iceland, Liechtenstein and Norway registers. | -UK Companies House deserves a specific mention: it is free, has a clean API, and publishes full -filing history including scanned documents. It is the best-documented major registry. +## Key tools + +### Companies House API + +The reference implementation of what a company registry ought to be: free, keyed, documented, +and complete back to incorporation including the scanned paper filings. Use it as the model for how +much you should expect elsewhere, and be disappointed accordingly. + +Register at +[developer.company-information.service.gov.uk](https://developer.company-information.service.gov.uk/) +for a free key. Authentication is HTTP basic with the key as the username and an empty password — +note the trailing colon. ```bash -# Companies House API — free key from developer.company-information.service.gov.uk +export CH_KEY='your-key-here' +``` + +```bash +# find the company and get its number, which is what everything else keys on +curl -s -u "$CH_KEY:" \ + 'https://api.company-information.service.gov.uk/search/companies?q=example+trading+ltd' \ + | jq '.items[] | {title, company_number, company_status, address_snippet}' + +# the profile: incorporation date, SIC codes, registered office, accounts due +curl -s -u "$CH_KEY:" \ + 'https://api.company-information.service.gov.uk/company/12345678' \ + | jq '{company_name, date_of_creation, company_status, sic_codes, registered_office_address}' + +# officers, including resignation dates — the sideways pivot starts here +curl -s -u "$CH_KEY:" \ + 'https://api.company-information.service.gov.uk/company/12345678/officers' \ + | jq '.items[] | {name, officer_role, appointed_on, resigned_on, links}' + +# persons with significant control: the UK's beneficial-ownership disclosure +curl -s -u "$CH_KEY:" \ + 'https://api.company-information.service.gov.uk/company/12345678/persons-with-significant-control' \ + | jq '.items[] | {name, kind, natures_of_control, notified_on, ceased_on}' + +# every other appointment this officer holds — nominee detection in one call +curl -s -u "$CH_KEY:" \ + 'https://api.company-information.service.gov.uk/officers/OFFICER_ID/appointments' \ + | jq '{total:.total_results, items:[.items[] | .appointed_to.company_name]}' + +# charges and mortgages, which name lenders the company never advertised +curl -s -u "$CH_KEY:" \ + 'https://api.company-information.service.gov.uk/company/12345678/charges' \ + | jq '.items[] | {created_on, status, persons_entitled}' + +# the filing history, then fetch the scanned document itself curl -s -u "$CH_KEY:" \ - 'https://api.company-information.service.gov.uk/search/companies?q=example+ltd' | jq '.items[] | {title, company_number, company_status}' + 'https://api.company-information.service.gov.uk/company/12345678/filing-history' \ + | jq '.items[] | {date, type, description, doc:.links.document_metadata}' +# search officers by name across all UK companies curl -s -u "$CH_KEY:" \ - 'https://api.company-information.service.gov.uk/company/12345678/officers' | jq '.items[] | {name, officer_role, appointed_on}' + 'https://api.company-information.service.gov.uk/search/officers?q=jane+smith' \ + | jq '.items[] | {title, appointment_count, date_of_birth}' +``` + +Rate limit is 600 requests per five minutes per key; exceed it and you get HTTP 429 with a +`Retry-After`. The document-download host is `document-api.company-information.service.gov.uk`, not +the API host, and it needs the same basic auth plus an `Accept: application/pdf` header. + +What it does not tell you: a PSC statement is self-declared and unverified, the 25% threshold means +real control can sit legally just underneath disclosure, and `date_of_birth` is published as month +and year only. "PSC: none" sometimes means genuinely dispersed ownership and sometimes means nobody +filed. + +### OpenCorporates + +The aggregator that normalises ~200 jurisdictions into one schema, which is what makes it the right +first call when you do not yet know where a company is registered. Its value is breadth and +cross-jurisdiction name matching; its weakness is freshness, so confirm anything important against +the registry it came from. + +**The API is key-gated and mostly paid.** Free keys exist for open-data projects that republish +under a share-alike licence, and journalists, academics and NGOs can apply for free or discounted +access, but the free allowance is small — of the order of 50 requests a day. Commercial plans start +in the low thousands of pounds a year. Without a key the API returns HTTP 401; the website search +remains free. + +```bash +export OC_TOKEN='your-api-token' +``` + +```bash +# name search across every jurisdiction OpenCorporates holds +curl -s -G 'https://api.opencorporates.com/v0.4/companies/search' \ + --data-urlencode 'q=example trading' --data-urlencode "api_token=$OC_TOKEN" \ + | jq '.results.companies[].company | {name, jurisdiction_code, company_number, current_status}' + +# narrow to one jurisdiction once you know it +curl -s -G 'https://api.opencorporates.com/v0.4/companies/search' \ + --data-urlencode 'q=example trading' --data-urlencode 'jurisdiction_code=gb' \ + --data-urlencode "api_token=$OC_TOKEN" | jq '.results.total_count' + +# the canonical record, including the source URL you should cite instead +curl -s -G 'https://api.opencorporates.com/v0.4/companies/gb/12345678' \ + --data-urlencode "api_token=$OC_TOKEN" \ + | jq '.results.company | {name, incorporation_date, registered_address_in_full, source}' + +# officers attached to that company +curl -s -G 'https://api.opencorporates.com/v0.4/companies/gb/12345678' \ + --data-urlencode "api_token=$OC_TOKEN" \ + | jq '.results.company.officers[].officer | {name, position, start_date, end_date}' + +# search officers by name globally — the cross-border nominee sweep +curl -s -G 'https://api.opencorporates.com/v0.4/officers/search' \ + --data-urlencode 'q=jane smith' --data-urlencode "api_token=$OC_TOKEN" \ + | jq '.results.officers[].officer | {name, jurisdiction_code, company:.company.name}' + +# every company at one address, which is how you find a formation agent +curl -s -G 'https://api.opencorporates.com/v0.4/companies/search' \ + --data-urlencode 'q=' --data-urlencode 'registered_address=*Tortola*' \ + --data-urlencode "api_token=$OC_TOKEN" | jq '.results.total_count' + +# check your remaining allowance before a loop burns the day's quota +curl -s -G 'https://api.opencorporates.com/v0.4/account_status' \ + --data-urlencode "api_token=$OC_TOKEN" | jq '.results.account_status.usage' +``` + +Every company record carries a `source` object with the registry URL and the date OpenCorporates +scraped it. That date is the one thing you must read: a record last refreshed two years ago will +show a director who resigned eighteen months back as current. + +### SEC EDGAR full-text search + +Every US filing since 2001, searchable by phrase. This is the one free full-text corporate archive +of real size, and it is the fastest way to find a private person or entity named in someone else's +filing — a subsidiary list, a related-party note, a beneficial-ownership schedule. + +There is no API key. The identification mechanism is the `User-Agent` header, which must name your +application and a contact address; send a generic one and you get HTTP 403 and a short IP block. +The limit is ten requests per second across all `sec.gov` hosts. + +```bash +UA='DAEMON-SEC research you@example.com' +``` + +```bash +# phrase search across all filings — quotes force an exact phrase +curl -s -A "$UA" -G 'https://efts.sec.gov/LATEST/search-index' \ + --data-urlencode 'q="Example Trading Limited"' \ + | jq '{total:.hits.total.value, hits:[.hits.hits[] | {name:._source.display_names, date:._source.file_date, id:._id}]}' + +# restrict to a form type: beneficial-ownership schedules +curl -s -A "$UA" -G 'https://efts.sec.gov/LATEST/search-index' \ + --data-urlencode 'q="Example Trading"' --data-urlencode 'forms=SC 13D' \ + | jq '.hits.total.value' + +# restrict to a date window — dateRange=custom is required alongside the dates +curl -s -A "$UA" -G 'https://efts.sec.gov/LATEST/search-index' \ + --data-urlencode 'q="going concern"' --data-urlencode 'forms=10-K' \ + --data-urlencode 'dateRange=custom' \ + --data-urlencode 'startdt=2025-01-01' --data-urlencode 'enddt=2025-06-30' \ + | jq '.hits.total.value' + +# everything one filer said about a term, by CIK +curl -s -A "$UA" -G 'https://efts.sec.gov/LATEST/search-index' \ + --data-urlencode 'q="related party"' --data-urlencode 'entityName=0000320193' \ + | jq '.hits.total.value' + +# page past the first ten results (the corpus caps at 10,000 total) +curl -s -A "$UA" -G 'https://efts.sec.gov/LATEST/search-index' \ + --data-urlencode 'q="Example Trading"' --data-urlencode 'from=10' \ + | jq '[.hits.hits[]._id]' + +# a filer's complete submission index, from the structured-data host +curl -s -A "$UA" 'https://data.sec.gov/submissions/CIK0000320193.json' \ + | jq '{name, cik, formerNames, tickers, recent:[.filings.recent.form[0:5]]}' + +# resolve a ticker to a CIK, which every other endpoint wants zero-padded +curl -s -A "$UA" 'https://www.sec.gov/files/company_tickers.json' \ + | jq -r 'to_entries[] | select(.value.ticker=="AAPL") | .value.cik_str' + +# one XBRL concept across a company's whole history +curl -s -A "$UA" \ + 'https://data.sec.gov/api/xbrl/companyconcept/CIK0000320193/us-gaap/Revenues.json' \ + | jq '.units["USD"] | length' +``` + +The `_id` in a hit is `accession-number:filename`, which reconstructs to a document URL under +`https://www.sec.gov/Archives/edgar/data/<cik>/<accession-no-dashes>/<filename>`. The result set is +capped at 10,000 and paged ten at a time, so a broad query needs narrowing by form and date rather +than paging. + +Full-text coverage starts in 2001. Anything older is in EDGAR but only findable by filer and form, +not by phrase — which matters when the structure you are chasing was set up in the 1990s. + +### edgartools + +A Python layer over EDGAR that parses filings into typed objects rather than handing you HTML, and +handles the `User-Agent` requirement and rate limiting for you. Use it the moment you need more +than a handful of filings, and particularly for Form 4 insider transactions and 13F holdings, where +the raw XML is tedious and the parsing is where the mistakes happen. + +```bash +pip install edgartools +``` + +```python +from edgar import Company, set_identity, get_filings + +set_identity("you@example.com") # required; this becomes the User-Agent + +c = Company("AAPL") + +# the filing index for one form type +c.get_filings(form="10-K").head(5) + +# insider transactions as structured objects, not XML +c.get_filings(form="4").latest().obj() + +# institutional holdings from a 13F, already parsed into positions +get_filings(form="13F-HR").latest().obj().holdings + +# subsidiary and related-party text out of the latest annual report +tenk = c.get_filings(form="10-K").latest().obj() +print(tenk["Item 1"][:2000]) + +# one XBRL concept as a DataFrame, for a time series +c.get_facts().query().by_concept("Revenues").to_dataframe() + +# filings across all companies in a quarter, to sweep a form type +get_filings(year=2025, quarter=1, form="SC 13D") ``` -EDGAR full-text search is equally scriptable: +`set_identity` is not optional — without it the library refuses to call EDGAR, which is the correct +behaviour. For bulk retrieval of raw documents rather than parsed objects, +`sec-edgar-downloader` is the simpler tool: `Downloader("YourOrg", "you@example.com")` then +`dl.get("10-K", "AAPL", after="2020-01-01", limit=5)`. + +### OpenSanctions + +Consolidated sanctions designations, politically-exposed-person lists and persons of interest from +several hundred sources, normalised into one entity model. Run every name you extract from a +registry through it, because a designation changes the legal meaning of the whole structure. + +The API is metered pay-as-you-go, with **free keys available on request for journalism, +civil-society and academic work**. The full dataset is also published for download under a +share-alike licence, and the matching service is open source as +[yente](https://github.com/opensanctions/yente), which you can self-host with no key at all — the +right choice when the names you are screening should not leave your machine. + +```bash +export OS_API_KEY='your-key' +``` + +```bash +# free-text search across the default collection +curl -s -H "Authorization: ApiKey $OS_API_KEY" -G \ + 'https://api.opensanctions.org/search/default' --data-urlencode 'q=Ilham Aliyev' \ + | jq '.results[] | {id, caption, schema, datasets, topics}' + +# restrict to companies rather than people +curl -s -H "Authorization: ApiKey $OS_API_KEY" -G \ + 'https://api.opensanctions.org/search/default' \ + --data-urlencode 'q=Example Trading' --data-urlencode 'schema=Company' \ + | jq '.total.value' + +# structured matching, which scores candidates instead of string-matching +curl -s -H "Authorization: ApiKey $OS_API_KEY" -X POST \ + -H 'Content-Type: application/json' \ + 'https://api.opensanctions.org/match/default' \ + -d '{"queries":{"q1":{"schema":"Person","properties":{"name":["Jane Smith"],"birthDate":["1970"],"nationality":["gb"]}}}}' \ + | jq '.responses.q1.results[] | {score, caption, topics}' + +# the full entity, including every alias, address and linked entity +curl -s -H "Authorization: ApiKey $OS_API_KEY" \ + 'https://api.opensanctions.org/entities/NK-ENTITY-ID' | jq '.properties' + +# what datasets a hit came from, which is how you cite the designation +curl -s -H "Authorization: ApiKey $OS_API_KEY" \ + 'https://api.opensanctions.org/datasets' | jq '.datasets[] | {name, title, entity_count}' + +# self-hosted yente takes the same paths with no key +curl -s -G 'http://localhost:8000/search/default' --data-urlencode 'q=Jane Smith' | jq '.total' +``` + +A match is a match on *data*, not an identification. Common names produce confident-looking hits on +unrelated people, and the sanctions lists themselves contain transliteration variants of one person +as separate entries. Always open the entity, read the aliases and birth date, and cite the +underlying designation — the OFAC or EU legal act — not the aggregator. + +### OCCRP Aleph + +Registries, leaks, court records, procurement data and gazettes in one index, all mapped onto the +FollowTheMoney entity schema. That shared schema is the point: a `Company` from a Latvian registry +and a `Company` from a leaked corporate database have the same property names, so you can ask one +query across both. + +API keys come with an Aleph account; some collections need separate permission and investigative +collections are not public. The header format is `Authorization: ApiKey <key>`, and all paths sit +under `/api/2/`. ```bash -curl -s 'https://efts.sec.gov/LATEST/search-index?q=%22specific+phrase%22&forms=10-K' \ - -H 'User-Agent: researcher you@example.com' | jq '.hits.hits[]._source | {display_names, file_date}' +export ALEPH_KEY='your-api-key' +H="Authorization: ApiKey $ALEPH_KEY" ``` -EDGAR requires a `User-Agent` identifying you, and will block you if you omit it. +```bash +# free-text search across every collection you can see +curl -s -H "$H" -G 'https://aleph.occrp.org/api/2/entities' \ + --data-urlencode 'q=Example Trading Limited' \ + | jq '.results[] | {id, schema, caption, collection:.collection.label}' + +# restrict by entity type — the schema names are the FollowTheMoney ones +curl -s -H "$H" -G 'https://aleph.occrp.org/api/2/entities' \ + --data-urlencode 'q=Jane Smith' --data-urlencode 'filter:schema=Person' \ + | jq '.total' + +# restrict to one collection, once you know which dataset matters +curl -s -H "$H" -G 'https://aleph.occrp.org/api/2/entities' \ + --data-urlencode 'q=Example Trading' --data-urlencode 'filter:collection_id=123' \ + | jq '.total' + +# one entity in full, with every property the schema allows +curl -s -H "$H" 'https://aleph.occrp.org/api/2/entities/ENTITY_ID' | jq '.properties' + +# what this entity is connected to and how — ownership, directorship, payment +curl -s -H "$H" -G 'https://aleph.occrp.org/api/2/entities/ENTITY_ID/expand' \ + --data-urlencode 'limit=50' \ + | jq '.results[] | {property, count, entities:[.entities[].caption]}' + +# the entity-model definitions, so you know which properties exist to filter on +curl -s 'https://aleph.occrp.org/api/2/metadata' \ + | jq '.schemata | keys | map(select(. | test("Company|Ownership|Person|Directorship")))' + +# list the collections your key can reach, with their ids and update dates +curl -s -H "$H" -G 'https://aleph.occrp.org/api/2/collections' --data-urlencode 'limit=50' \ + | jq '.results[] | {id, label, category, updated_at}' +``` + +The schema is what makes Aleph worth the learning curve. `Company` and `Person` are the nodes; +`Ownership`, `Directorship`, `Membership` and `Payment` are *intermediate entities* with their own +properties — an `Ownership` carries a percentage, a start date and an end date, and connects an +owner to an asset. `/expand` walks those edges, which is how you traverse a chain without reading +every document. + +Leaked collections are snapshots. An `Ownership` from a 2016 leak describes 2016, and saying so is +the difference between a finding and a libel risk. + +### ICIJ Offshore Leaks + +Entities, officers, intermediaries and addresses from the Panama, Paradise, Pandora, Bahamas and +Offshore Leaks investigations. Web-only, no API, and worth a walkthrough because the structure of +the database is not obvious from the search box. + +Search at [offshoreleaks.icij.org](https://offshoreleaks.icij.org/) by company name, person name or +address. Read the result in this order. + +```text +Entity the offshore company itself: jurisdiction, incorporation and inactivation + dates, and the leak it came from. The jurisdiction plus the dates is what + you take back to that jurisdiction's registry +Officers the people and companies attached, each with a role (shareholder, director, + beneficiary, nominee). "Nominee" here is explicit, which is rarer than it + sounds +Intermediary the law firm or agent that formed it. This is the strongest pivot in the + database — the intermediary's other clients are listed, and a pattern of + one agent forming a cluster of companies is a structure +Addresses search by address, not name, when a name is too common. A residential + address shared by forty companies is the finding +Linked the graph view, which shows the chain between two entities if one exists +``` + +Capture the node URL for every entity and officer, the leak name, and the date range the data +covers. The URLs are stable and citable. + +Two limits to state every time you use it: being in the database is not evidence of wrongdoing — +offshore structures are legal and ICIJ says so prominently — and the data is as of the leak, which +for Panama means 2015. A company shown as active was active then. + +### Jurisdiction-by-jurisdiction fallbacks + +When the chain crosses a border, what you can get changes completely. These are the sources worth +knowing before you conclude a jurisdiction publishes nothing. + +| Source | What it adds | +| --- | --- | +| [BRIS / e-Justice](https://e-justice.europa.eu/content_find_a_company-489-en.do) | One search box over most EU registers. Coverage per country varies from full filings to name-and-number only. | +| [North Data](https://northdata.com) | EU registers with relationship graphs and German trade-register notices, which often name shareholders the registry does not. | +| [Open Ownership Register](https://register.openownership.org/) | Beneficial-ownership declarations consolidated across the countries that publish them. | +| [RuPEP](https://rupep.org/en/) | Politically exposed persons in Russia, Belarus, Kazakhstan, Kyrgyzstan, Georgia and Moldova, with family and associate links. | +| [SanctionsExplorer](https://sanctionsexplorer.org/) | Historical OFAC, UN and EU designations, including delisted ones — which is what you need for a structure that predates a current list. | +| [ImportYeti](https://www.importyeti.com/) | US customs sea-shipment records: who ships to whom, which evidences a trading relationship no registry records. | +| [Wikipedia's register list](https://en.wikipedia.org/wiki/List_of_official_business_registers) | The maintained index of official registers worldwide. Start here for a jurisdiction you have not worked before. | ## Reading the structure - **Nominee directors** appear on dozens or hundreds of companies. A director with 200 appointments - is a service provider, not a decision-maker. -- **Registered-address clustering** — many companies at one address usually means a formation agent. + is a service provider, not a decision-maker, and the Companies House appointments endpoint tells + you which you have in one call. +- **Registered-address clustering** usually means a formation agent. Search the address, not the + name, and the agent's whole book appears. - **Shareholders that are themselves companies** are the chain you have to walk. Keep going until - you hit a natural person or a jurisdiction that will not tell you. + you hit a natural person or a jurisdiction that will not tell you, then record which. - **Charges and mortgages** name lenders, which reveals banking relationships the company did not - advertise. + advertise and gives you a regulated counterparty who had to do their own due diligence. +- **Dates are the load-bearing part.** An ownership that ended before the event you are + investigating is not relevant, and a leak snapshot has one date for everything in it. ## Tool reference @@ -122,6 +507,95 @@ EDGAR requires a `User-Agent` identifying you, and will block you if you omit it - **Dissolved does not mean gone.** Historical filings often hold what you need, and some registries purge them after a few years — capture early. +## Worked example + +One datum: the name "Northgate Minerals Trading" on an invoice. The company, the numbers and the +returns are invented; the order of the pivots and what each source can and cannot settle are not. + +```bash +# 1. where does it exist at all +curl -s -G 'https://api.opencorporates.com/v0.4/companies/search' \ + --data-urlencode 'q=Northgate Minerals Trading' --data-urlencode "api_token=$OC_TOKEN" \ + | jq '.results.companies[].company | {name, jurisdiction_code, company_number, current_status}' +# two hits: gb/09876543 (active) and vg/1654321 (OpenCorporates record, source dated 2019) +``` + +A UK company and a British Virgin Islands company with the same name. The UK one is where the data +is, so start there and carry the number `09876543`. + +```bash +# 2. who is declared to control it +curl -s -u "$CH_KEY:" \ + 'https://api.company-information.service.gov.uk/company/09876543/persons-with-significant-control' \ + | jq '.items[] | {name, kind, natures_of_control}' +# one PSC, kind "corporate-entity-person-with-significant-control", +# name "Northgate Holdings Ltd", country of registration "Virgin Islands, British" +``` + +The declared controller is the BVI company, so UK disclosure has handed the chain straight back +offshore. That is the expected outcome, and worth stating as one. + +```bash +# 3. the officers, and whether they are real or nominal +curl -s -u "$CH_KEY:" \ + 'https://api.company-information.service.gov.uk/company/09876543/officers' \ + | jq '.items[] | {name, officer_role, appointed_on, id:.links.officer.appointments}' +# two directors; take each appointments link +curl -s -u "$CH_KEY:" \ + 'https://api.company-information.service.gov.uk/officers/AbC123.../appointments' | jq '.total_results' +# 147 +``` + +147 appointments makes that director a formation-agent nominee, not a decision-maker. The other has +three, all in the same group — that is the person to follow. + +```bash +# 4. screen both names and the BVI entity +curl -s -H "Authorization: ApiKey $OS_API_KEY" -G \ + 'https://api.opensanctions.org/search/default' --data-urlencode 'q=Northgate Holdings' \ + | jq '.results[] | {caption, schema, topics, datasets}' +# no designation; one PEP-adjacent hit on a similarly-named Cyprus entity — not the same company +``` + +A near-miss on a different entity is exactly the false positive to discard explicitly, in writing, +so it does not resurface later as a claim. + +```bash +# 5. the offshore end, where the registry publishes nothing +curl -s -H "Authorization: ApiKey $ALEPH_KEY" -G 'https://aleph.occrp.org/api/2/entities' \ + --data-urlencode 'q=Northgate Holdings' --data-urlencode 'filter:schema=Company' \ + | jq '.results[] | {id, caption, collection:.collection.label}' +# a hit in an ICIJ collection; expand it for the ownership edges +curl -s -H "Authorization: ApiKey $ALEPH_KEY" -G \ + 'https://aleph.occrp.org/api/2/entities/ENTITY_ID/expand' --data-urlencode 'limit=50' \ + | jq '.results[] | {property, entities:[.entities[].caption]}' +# Ownership -> a named individual, percentage 100, startDate 2014 +``` + +Cross-checking that individual in [ICIJ Offshore Leaks](https://offshoreleaks.icij.org/) gives the +same person as beneficiary via the same BVI intermediary, with a residential address shared by +eleven other companies — a second, independent confirmation of the same edge. + +```bash +# 6. does the name appear in anyone's US filings +curl -s -A "$UA" -G 'https://efts.sec.gov/LATEST/search-index' \ + --data-urlencode 'q="Northgate Minerals Trading"' \ + | jq '.hits.hits[] | {name:._source.display_names, date:._source.file_date}' +# one 10-K exhibit, 2023, listing it as a counterparty of a listed miner +``` + +The defensible chain: invoice name → UK company 09876543 (Companies House, retrieved with date) → +corporate PSC in BVI (UK filing, self-declared) → named individual with 100% ownership from 2014 +(Aleph, ICIJ collection, snapshot date stated) → confirmed independently in the ICIJ public +database via the same intermediary → named as a counterparty in a 2023 SEC exhibit. Six sources, +two of them independent of each other on the key link. + +What you still cannot say: whether that individual controls it *today*. The ownership edge is from +a leak snapshot, the BVI publishes nothing current, and the UK PSC declaration names the company +rather than the person. "Beneficial ownership is not currently published in the British Virgin +Islands; the most recent evidenced owner is X as of 2014" is the honest sentence, and it is a +stronger finding than an unqualified assertion. + ## Broader catalogues - [Public Records OSINT](https://tools.osintnewsletter.com/tool-categories/public-records-osint) diff --git a/src/content/sheets/osint/conflict-and-environment.md b/src/content/sheets/osint/conflict-and-environment.md @@ -4,9 +4,9 @@ description: "Event datasets, munitions identification, fire and deforestation f category: osint subcategory: "Thematic" tags: [osint, conflict, environment, monitoring, satellite] -tools: [acled, liveuamap, global-forest-watch, firms] +tools: [acled, gdelt, firms, sentinel-hub, global-forest-watch, osmp, liveuamap] difficulty: intermediate -updated: 2026-09-28 +updated: 2026-10-03 references: - name: "Bellingcat's Online Investigation Toolkit" url: "https://bellingcat.gitbook.io/toolkit" @@ -25,54 +25,460 @@ references: ## What this covers Two thematic areas with mature, purpose-built open datasets: armed conflict and environmental -change. Both are unusual in OSINT for having authoritative structured data rather than just tools — -which means the skill is in reading the datasets' methodology, not in finding them. +change. Both are unusual in OSINT for having authoritative structured data behind them rather than +just tools, which moves the skill from finding sources to reading methodologies and querying them +properly. + +## Method + +1. **Start from a bounded question with a place and a date range.** These are datasets, not search + engines — a query without a bounding box and a window either returns nothing useful or exceeds + your quota. +2. **Read the codebook before the data.** What ACLED counts as an event, how FIRMS defines a + thermal anomaly, what a Global Forest Watch alert confidence class means: each materially + changes what a number means, and each is documented. +3. **Query the dataset rather than the dashboard.** Dashboards aggregate in ways you cannot audit + and cannot cite. Pull the records, store them with the retrieval date, and do your own + aggregation. +4. **Establish the baseline before the anomaly.** One fire detection at an industrial site means + nothing until you know how often that pixel normally burns. Pull the same query for the previous + year and the previous month. +5. **Corroborate across two independent feeds.** Independence means different sensors or different + collection methods — a VIIRS hotspot and a Sentinel-2 burn scar, an ACLED event and a dated + photograph, a GDELT article cluster and a satellite image. Two aggregators that both scrape the + same wire service are one source. +6. **For munitions, work from markings outward**, and write "consistent with" unless you have + stencilling, lot numbers or a fuze you can read. This is the single most common place + open-source conflict reporting goes wrong. +7. **Record the evidentiary chain for each claim**: the dataset and its version, the exact query, + the retrieval timestamp, the hash of any imagery you downloaded, and the second source. A + finding that cannot be re-run is not a finding. ## Conflict | Source | What it gives you | | --- | --- | -| [ACLED](https://acleddata.com/) | Coded conflict events — date, location, actors, fatalities — globally, with a documented methodology and a free API. The reference dataset. | -| [Liveuamap](https://liveuamap.com/) | Near-real-time mapped events aggregated from social media. Fast but unverified; treat as a lead generator. | +| [ACLED](https://acleddata.com/) | Coded conflict events — date, location, actors, event type, fatalities — globally, with a published methodology and an API. The reference dataset. | +| [GDELT](https://www.gdeltproject.org/) | Global news coverage indexed by location, theme, tone and actor, updated every 15 minutes. Not events; coverage of events. | +| [Liveuamap](https://liveuamap.com/) | Near-real-time mapped events aggregated from social media. Fast but unverified; a lead generator only. | +| [Geoconfirmed](https://geoconfirmed.org/) | Volunteer-geolocated conflict media, each entry with the coordinates and the reasoning. Verified work you can audit. | | [Open Source Munitions Portal](https://osmp.ngo/) | Identification reference for munitions and their remnants, which is how you turn a photo of debris into a weapon type. | -| [Bellingcat's Civilian Harm datasets](https://ukraine.bellingcat.com/) | Verified incident databases with sourcing for each entry. | +| [Bellingcat's Civilian Harm datasets](https://ukraine.bellingcat.com/) | Verified incident databases with the sourcing attached to each entry. | + +### ACLED API + +Coded event data with a documented methodology, which is what separates it from every aggregator in +the table above. Each record carries a date, coordinates with a stated precision level, named +actors, an event type and sub-event type, a fatality estimate, and the sources it was coded from. -ACLED's API is the one to build on: +**The API moved and the auth model changed.** The old `api.acleddata.com` host with +`key=`/`email=` query parameters is gone — the hostname no longer resolves, and the key system was +retired in September 2025. Current access is a bearer token from an OAuth password grant against +your myACLED account, used against `acleddata.com/api/`. ```bash -curl -s 'https://api.acleddata.com/acled/read?key=KEY&email=you@example.com&country=Ukraine&event_date=2024-01-01|2024-03-31&event_date_where=BETWEEN&limit=0' \ - | jq '.data | length' +# token: valid 24 hours, with a refresh token valid 14 days +TOKEN=$(curl -s -X POST 'https://acleddata.com/oauth/token' \ + -d 'username=you@example.com' --data-urlencode 'password=YOUR_PASSWORD' \ + -d 'grant_type=password' -d 'client_id=acled' -d 'scope=authenticated' \ + | jq -r .access_token) ``` -Read ACLED's codebook before drawing conclusions from it. What counts as an "event", how fatalities -are attributed, and what sourcing threshold applies all materially affect the numbers, and -comparisons across regions can be distorted by differences in local reporting density. +```bash +# events in one country over a window, as JSON +curl -s -H "Authorization: Bearer $TOKEN" -G 'https://acleddata.com/api/acled/read' \ + --data-urlencode 'country=Ukraine' \ + --data-urlencode 'event_date=2026-01-01|2026-03-31' \ + --data-urlencode 'event_date_where=BETWEEN' --data-urlencode 'limit=0' \ + | jq '.count' -**Munitions identification** is a specialist skill and the single most common place where -open-source conflict reporting goes wrong. Match on markings, dimensions and fragmentation pattern -against a reference, and say "consistent with" rather than "is" unless you have markings. +# CSV straight to disk, which is what you archive +curl -s -H "Authorization: Bearer $TOKEN" -G 'https://acleddata.com/api/acled/read' \ + --data-urlencode '_format=csv' --data-urlencode 'country=Sudan' \ + --data-urlencode 'event_date=2026-01-01|2026-03-31' \ + --data-urlencode 'event_date_where=BETWEEN' --data-urlencode 'limit=0' \ + -o "acled-sudan-$(date -u +%F).csv" -## Environment +# one event type only — the sub-event taxonomy is where the precision is +curl -s -H "Authorization: Bearer $TOKEN" -G 'https://acleddata.com/api/acled/read' \ + --data-urlencode 'country=Nigeria' \ + --data-urlencode 'event_type=Explosions/Remote violence' \ + --data-urlencode 'event_date=2026-01-01|2026-03-31' \ + --data-urlencode 'event_date_where=BETWEEN' | jq '.data | length' -| Source | What it gives you | -| --- | --- | -| [Global Forest Watch](https://www.globalforestwatch.org/) | Deforestation alerts and tree-cover change, near-real-time. | -| [NASA FIRMS](https://firms.modaps.eosdis.nasa.gov/) | Active fire and thermal anomaly detections, updated several times a day. | -| [Global Fishing Watch](https://globalfishingwatch.org/map) | Apparent fishing effort inferred from AIS; exposes probable illegal fishing. | -| [Sentinel Hub](https://apps.sentinel-hub.com/eo-browser/) | Free Sentinel-2 imagery with band combinations for burn scars, water and vegetation. | +# a named actor's events, for an order-of-battle or attribution question +curl -s -H "Authorization: Bearer $TOKEN" -G 'https://acleddata.com/api/acled/read' \ + --data-urlencode 'actor1=Military Forces of Sudan (2019-)' \ + --data-urlencode 'actor1_where=%3D' --data-urlencode 'limit=200' | jq '.count' + +# narrow to an admin region, which is how you bound to a locality +curl -s -H "Authorization: Bearer $TOKEN" -G 'https://acleddata.com/api/acled/read' \ + --data-urlencode 'country=Mali' --data-urlencode 'admin1=Mopti' \ + --data-urlencode 'event_date=2026-01-01|2026-03-31' \ + --data-urlencode 'event_date_where=BETWEEN' | jq '.data | length' + +# only the fields you need, which keeps responses small +curl -s -H "Authorization: Bearer $TOKEN" -G 'https://acleddata.com/api/acled/read' \ + --data-urlencode 'country=Ukraine' --data-urlencode 'limit=50' \ + --data-urlencode 'fields=event_date|event_type|sub_event_type|location|latitude|longitude|fatalities|source' \ + | jq '.data[0]' + +# the deleted-records endpoint, so a stored copy can be reconciled with upstream revisions +curl -s -H "Authorization: Bearer $TOKEN" -G 'https://acleddata.com/api/deleted/read' \ + --data-urlencode 'deleted_timestamp=1735689600' | jq '.count' +``` + +The `R` package `acledR` and the `acled` Python package both wrap this and handle the token +refresh; either is less work than scripting the grant yourself if you query regularly. Access is +free for non-commercial use with registration, and the terms restrict redistribution of the raw +data — cite and link rather than re-publishing the CSV. + +Read the codebook before drawing any conclusion. Three things in particular: coordinate precision +is a graded field, and a precision-3 event is located to an admin region rather than a point; +fatality figures are the most conservative reported estimate, not a count; and event density tracks +reporting density, so comparing two regions compares their journalism as much as their violence. +The `deleted` endpoint exists because ACLED revises history as better sourcing arrives, which means +a stored extract drifts from upstream. + +### GDELT DOC 2.0 API + +Global news coverage, indexed every 15 minutes across languages, queryable by phrase, location, +theme, tone and source country. It tells you what was *reported*, where and in what volume, which +is a different and complementary question to what happened. No key, no registration. + +```bash +# articles matching a phrase, geographically filtered, as JSON +curl -s -G 'https://api.gdeltproject.org/api/v2/doc/doc' \ + --data-urlencode 'query="artillery" sourcecountry:UP' \ + --data-urlencode 'mode=artlist' --data-urlencode 'format=json' \ + --data-urlencode 'maxrecords=100' --data-urlencode 'timespan=7d' \ + | jq '.articles[] | {seendate, domain, title, language}' + +# coverage volume over time — the shape of the curve is the finding +curl -s -G 'https://api.gdeltproject.org/api/v2/doc/doc' \ + --data-urlencode 'query="port of berbera"' \ + --data-urlencode 'mode=timelinevol' --data-urlencode 'format=json' \ + --data-urlencode 'timespan=3m' | jq '.timeline[0].data[-10:]' + +# an exact window rather than a rolling span +curl -s -G 'https://api.gdeltproject.org/api/v2/doc/doc' \ + --data-urlencode 'query="chemical plant" sourcelang:eng' \ + --data-urlencode 'mode=artlist' --data-urlencode 'format=json' \ + --data-urlencode 'startdatetime=20260301000000' \ + --data-urlencode 'enddatetime=20260315000000' | jq '.articles | length' + +# which countries are covering a story, and which are not +curl -s -G 'https://api.gdeltproject.org/api/v2/doc/doc' \ + --data-urlencode 'query="grain corridor"' \ + --data-urlencode 'mode=timelinesourcecountry' --data-urlencode 'format=json' \ + --data-urlencode 'timespan=1m' | jq '[.timeline[].series] | .[0:8]' + +# tone distribution, for spotting a coordinated framing shift +curl -s -G 'https://api.gdeltproject.org/api/v2/doc/doc' \ + --data-urlencode 'query="sanctions relief"' \ + --data-urlencode 'mode=tonechart' --data-urlencode 'format=json' \ + --data-urlencode 'timespan=2w' | jq '.tonechart' + +# images published alongside the coverage, which sometimes surface the source photo +curl -s -G 'https://api.gdeltproject.org/api/v2/doc/doc' \ + --data-urlencode 'query="refinery fire"' \ + --data-urlencode 'mode=imagecollageinfo' --data-urlencode 'format=json' \ + --data-urlencode 'timespan=3d' | jq '.images[0:5]' +``` + +The DOC API serves a rolling window — roughly the last three months by default, reaching back to +1 January 2017 with explicit datetimes — and caps at 250 records per request for article and image +modes. For anything longer or larger, the GDELT Event and GKG tables are published on BigQuery and +as raw CSV, which is the right route for a multi-year series. + +What it does not tell you: GDELT indexes articles, not events, so one incident covered by four +hundred outlets is four hundred records. A volume spike measures attention. Syndication means the +same wire copy appears under many domains, which is why a GDELT cluster and an aggregator that +scrapes the same wires are not two independent sources. + +### NASA FIRMS + +Active fire and thermal anomaly detections from MODIS and VIIRS, available within three hours of +satellite overpass. Thermal detection is a far broader OSINT primitive than wildfire monitoring: +gas flaring, industrial process heat, shelling, burning buildings and crop residue all register. + +Get a free `MAP_KEY` from +[firms.modaps.eosdis.nasa.gov/api/map_key](https://firms.modaps.eosdis.nasa.gov/api/map_key/). -Thermal detections are a strong OSINT primitive well beyond wildfires: gas flaring, industrial -activity, shelling and burning all register. A FIRMS hotspot at an industrial site that should be -idle is a finding. +```bash +export MAP_KEY='your-map-key' +``` + +```bash +# check the key works and how much of your quota is left +curl -s "https://firms.modaps.eosdis.nasa.gov/mapserver/mapkey_status/?MAP_KEY=$MAP_KEY" + +# detections in a bounding box over the last day — order is west,south,east,north +curl -s "https://firms.modaps.eosdis.nasa.gov/api/area/csv/$MAP_KEY/VIIRS_SNPP_NRT/30.0,46.0,36.0,50.0/1" + +# a specific past date, with the day range starting from it (1-10 days) +curl -s "https://firms.modaps.eosdis.nasa.gov/api/area/csv/$MAP_KEY/VIIRS_NOAA20_NRT/30.0,46.0,36.0,50.0/2/2026-03-11" \ + -o firms-2026-03-11.csv + +# the archival standard-quality product, for anything older than the NRT window +curl -s "https://firms.modaps.eosdis.nasa.gov/api/area/csv/$MAP_KEY/VIIRS_SNPP_SP/30.0,46.0,36.0,50.0/5/2025-07-01" \ + -o firms-archive.csv + +# which products cover which dates, so you query one that exists +curl -s "https://firms.modaps.eosdis.nasa.gov/api/data_availability/csv/$MAP_KEY/all" + +# filter to high-confidence daytime detections at a site +curl -s "https://firms.modaps.eosdis.nasa.gov/api/area/csv/$MAP_KEY/VIIRS_SNPP_NRT/30.0,46.0,36.0,50.0/1" \ + | awk -F, 'NR==1 || ($10=="h" || $10>=80)' + +# a year of the same small box, to build the baseline before you call an anomaly +for d in 2025-0{1..9}-01 2025-1{0..2}-01; do + curl -s "https://firms.modaps.eosdis.nasa.gov/api/area/csv/$MAP_KEY/VIIRS_SNPP_SP/30.0,46.0,36.0,50.0/10/$d" + sleep 2 +done > baseline.csv +``` + +The quota is 5,000 transactions per 10-minute window per key, and a large bounding box counts as +several, so loop with a sleep. Day range per request is 1–10. VIIRS at 375 m resolves far more than +MODIS at 1 km and is what you want for anything at a single site; the `_NRT` products are the fast +ones, `_SP` the reprocessed archive. + +Reading a detection: `confidence` is `l`/`n`/`h` for VIIRS and a 0–100 percentage for MODIS, `frp` +is fire radiative power in megawatts, `daynight` matters because daytime detections suffer more +false positives from reflective surfaces, and `scan`/`track` give the actual pixel footprint — the +detection is a pixel, not a point, so the coordinates are the pixel centre and the real event is +somewhere within a few hundred metres of it. + +A hotspot is a thermal anomaly and nothing more. Gas flares burn continuously and appear every +single night, which is why the baseline query above matters: the finding is not "a hotspot at this +site", it is "a hotspot at a site that produced none in the preceding twelve months". + +### Sentinel Hub on the Copernicus Data Space Ecosystem -False-colour band combinations in EO Browser make change obvious that true colour hides: +Free Sentinel-1, Sentinel-2 and Sentinel-3 imagery, both as an interactive browser and as an API. +Use the browser to find the scene and settle the band combination; use the API to pull the exact +same thing reproducibly, with a date and a bounding box you can cite. + +**Access moved to the Copernicus Data Space Ecosystem**, which replaced both the old Copernicus +Open Access Hub and the commercial Sentinel Hub free tier for Sentinel data. Register at +[dataspace.copernicus.eu](https://dataspace.copernicus.eu/), create an OAuth client in your account +settings, and you get a monthly processing-unit quota at no cost. + +Start in [EO Browser](https://browser.dataspace.copernicus.eu/): set the area, set the date range, +pick Sentinel-2 L2A, and step through the available passes. The band combinations that make change +visible where true colour hides it: ```text -Burn scars (Sentinel-2): B12, B8A, B4 -Vegetation health (NDVI): (B8 - B4) / (B8 + B4) -Water / flooding: B8A, B11, B4 +Burn scars (Sentinel-2): B12, B8A, B4 — recent burns go red-brown, vegetation green +Vegetation health (NDVI): (B8 - B4) / (B8 + B4) +Normalised burn ratio (NBR): (B8 - B12) / (B8 + B12), differenced pre/post for severity +Water and flooding: B8A, B11, B4 — water goes near-black, wet soil distinct +Bare soil / earthworks: B11, B8, B2 — fresh excavation and berms separate from vegetation +Smoke and plumes: true colour, then B12/B8A/B4 to see through thin smoke +Sentinel-1 radar (SAR): VV/VH, for cloud cover, night, and detecting hulls at sea +``` + +Then reproduce it through the API: + +```bash +# token, valid about ten minutes +TOKEN=$(curl -s -X POST \ + 'https://identity.dataspace.copernicus.eu/auth/realms/CDSE/protocol/openid-connect/token' \ + -H 'content-type: application/x-www-form-urlencoded' \ + -d 'grant_type=client_credentials' -d 'client_id=YOUR_CLIENT_ID' \ + --data-urlencode 'client_secret=YOUR_CLIENT_SECRET' | jq -r .access_token) + +# which Sentinel-2 passes exist over this box in this window, and how cloudy +curl -s -X POST 'https://sh.dataspace.copernicus.eu/api/v1/catalog/1.0.0/search' \ + -H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \ + -d '{"collections":["sentinel-2-l2a"], + "bbox":[30.0,46.0,30.3,46.3], + "datetime":"2026-03-01T00:00:00Z/2026-03-20T00:00:00Z", + "limit":20, + "fields":{"include":["id","properties.datetime","properties.eo:cloud_cover"]}}' \ + | jq '.features[] | {id, datetime:.properties.datetime, cloud:.properties["eo:cloud_cover"]}' + +# render a false-colour burn-scar image for one date, as a GeoTIFF you can hash +curl -s -X POST 'https://sh.dataspace.copernicus.eu/api/v1/process' \ + -H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \ + -d '{"input":{"bounds":{"bbox":[30.0,46.0,30.3,46.3]}, + "data":[{"type":"sentinel-2-l2a", + "dataFilter":{"timeRange":{"from":"2026-03-11T00:00:00Z","to":"2026-03-12T00:00:00Z"}}}]}, + "output":{"width":1024,"height":1024,"responses":[{"identifier":"default","format":{"type":"image/tiff"}}]}, + "evalscript":"//VERSION=3\nfunction setup(){return {input:[\"B12\",\"B8A\",\"B04\"],output:{bands:3}}}\nfunction evaluatePixel(s){return [s.B12*2.5, s.B8A*2.5, s.B04*2.5]}"}' \ + -o burnscar-2026-03-11.tiff ``` +Hash the GeoTIFF the moment you download it and record the scene ID, the exact bounding box and the +acquisition datetime. That tuple is what makes the imagery claim reproducible — "a Sentinel-2 image +from around then" is not a citation. + +Quotas are monthly processing units and are replenished on the first of the month; a 1024-pixel +request is cheap, a full-resolution multi-band pull over a province is not. Sentinel-2 revisits +every five days at the equator and the optical bands see nothing through cloud, which is why +Sentinel-1 radar is the fallback for a specific date you cannot miss. + +### Global Forest Watch Data API + +Deforestation and fire alerts as a queryable dataset rather than a dashboard. The integrated alerts +layer combines three independent alert systems (GLAD-L, GLAD-S2 and RADD radar), which makes it one +of the few places where cross-sensor corroboration is already built in and labelled. + +Create an account, exchange it for a token, then mint an API key; the key goes in an `x-api-key` +header. Without one, every query returns a 401 telling you so. + +```bash +# sign up and get a short-lived token +curl -s -X POST 'https://data-api.globalforestwatch.org/auth/token' \ + -d 'username=you@example.com' --data-urlencode 'password=YOUR_PASSWORD' | jq -r .data.access_token + +# mint an API key with that token +curl -s -X POST 'https://data-api.globalforestwatch.org/auth/apikey' \ + -H "Authorization: Bearer $GFW_TOKEN" -H 'Content-Type: application/json' \ + -d '{"alias":"osint-research","organization":"research","email":"you@example.com","domains":[]}' \ + | jq -r .data.api_key +``` + +```bash +export GFW_KEY='your-api-key' + +# what datasets exist, and what each version is called +curl -s 'https://data-api.globalforestwatch.org/datasets' \ + | jq -r '.data[].dataset' | grep -E 'integrated_alerts|viirs|burned' + +# daily integrated deforestation alerts for one country +curl -s -H "x-api-key: $GFW_KEY" -G \ + 'https://data-api.globalforestwatch.org/dataset/gadm__integrated_alerts__iso_daily_alerts/latest/query/json' \ + --data-urlencode "sql=SELECT iso, gfw_integrated_alerts__date, SUM(alert__count) AS alerts + FROM results WHERE iso='BRA' AND gfw_integrated_alerts__date >= '2026-02-01' + GROUP BY iso, gfw_integrated_alerts__date ORDER BY 2 DESC" | jq '.data[0:5]' + +# down to a second-level administrative area, which is a usable locality +curl -s -H "x-api-key: $GFW_KEY" -G \ + 'https://data-api.globalforestwatch.org/dataset/gadm__integrated_alerts__adm2_daily_alerts/latest/query/json' \ + --data-urlencode "sql=SELECT adm1, adm2, SUM(alert__count) AS alerts FROM results + WHERE iso='COL' AND gfw_integrated_alerts__date >= '2026-01-01' + GROUP BY adm1, adm2 ORDER BY alerts DESC LIMIT 20" | jq '.data[0:5]' + +# alert confidence, which is where the cross-sensor corroboration lives +curl -s -H "x-api-key: $GFW_KEY" -G \ + 'https://data-api.globalforestwatch.org/dataset/gadm__integrated_alerts__adm2_daily_alerts/latest/query/json' \ + --data-urlencode "sql=SELECT gfw_integrated_alerts__confidence, SUM(alert__count) FROM results + WHERE iso='PER' GROUP BY 1" | jq '.data' + +# VIIRS fire alerts from the same API, so fire and forest loss share one query layer +curl -s -H "x-api-key: $GFW_KEY" -G \ + 'https://data-api.globalforestwatch.org/dataset/gadm__viirs__adm2_daily_alerts/latest/query/json' \ + --data-urlencode "sql=SELECT adm1, adm2, SUM(alert__count) FROM results + WHERE iso='IDN' AND alert__date >= '2026-02-01' GROUP BY 1,2 ORDER BY 3 DESC LIMIT 10" \ + | jq '.data[0:5]' + +# which fields a dataset actually has, before you guess a column name +curl -s -H "x-api-key: $GFW_KEY" \ + 'https://data-api.globalforestwatch.org/dataset/gadm__integrated_alerts__adm2_daily_alerts/latest/fields' \ + | jq -r '.data[].name' + +# CSV instead of JSON, for a spreadsheet handover +curl -s -H "x-api-key: $GFW_KEY" -G \ + 'https://data-api.globalforestwatch.org/dataset/gadm__integrated_alerts__iso_daily_alerts/latest/query/csv' \ + --data-urlencode "sql=SELECT * FROM results WHERE iso='BRA' LIMIT 1000" -o gfw-bra.csv +``` + +Versions are dated (`v20261003`), and `latest` resolves to the newest — pin the explicit version in +anything you cite, because `latest` moves under you. A `highest` confidence integrated alert means +more than one of the three systems detected the same loss, which is the built-in second source; a +`low` confidence alert is a single-sensor detection and should be treated as a lead. + +An alert is tree-cover loss, not a crime. Legal logging concessions, plantation harvest cycles, +storm damage and seasonal leaf-off all generate alerts. The question that turns an alert into a +story is whether the polygon falls inside a protected area or a concession boundary, and that is a +separate dataset join. + +### Open Source Munitions Portal + +A searchable reference library of verified photographs of munitions and their remnants, built by +Airwars and Armament Research Services. Web-only, no API, and the correct first stop when you have +a photograph of debris and no identification. Over a thousand verified entries, weighted heavily +towards the Middle East and Ukraine. + +Search at [osmp.ngo](https://osmp.ngo/) by munition category, country, date or free text, then work +the comparison deliberately: + +```text +1. Measure first find a scale reference in your photo — a hand, a boot, a kerb, a + standard brick — and estimate diameter and length before you + browse, so you are filtering rather than pattern-matching +2. Filter by category bomb, rocket, artillery projectile, guided missile, submunition. + Getting the category right eliminates most of the library +3. Compare the parts tail assembly and fin count, fuze well, lug spacing, nose profile, + body seams, driving band. These are discriminating; overall shape + is not +4. Read the stencilling alphanumeric markings, lot numbers, factory codes, colour bands. + A legible lot number is the only thing that gets you to a + production batch and a country of manufacture +5. Check the remnant fragmentation pattern and wall thickness distinguish families + that look identical intact +6. Cross-reference CAT-UXO for explosive ordnance detail and Bulletpicker's + manual archive for the original technical drawings +``` + +Record the OSMP entry ID for every comparison you made, including the ones you rejected, and the +specific features you matched on. "Consistent with a 9M27K rocket on the basis of fin count and +motor diameter, per OSMP entry NNNN" is defensible. "A cluster munition" is not, unless you can +read the markings. + +The geographic bias is a real limitation: a munition used outside the Middle East or Ukraine may +simply be absent from the library, and absence is not evidence of a different weapon. Pair OSMP +with [CAT-UXO](https://cat-uxo.com/) and +[Bulletpicker](https://bulletpicker.com/)'s scanned ordnance manuals, which have broader scope and +worse search. + +## Corroborating one event across two independent feeds + +The single most useful habit in this area, and the thing that makes a claim stand up. "Independent" +means a different sensor or a different collection route, not a different website. + +```text +Claim a munitions strike on a named facility, 11 March 2026, from a social post +Source A ACLED event: query country + admin1 + date, with event_type + Explosions/Remote violence. Returns a coded event, its own sources, a + coordinate precision level and a fatality estimate +Source B NASA FIRMS: VIIRS detections in a tight box around the facility for that + date and the day after. A thermal anomaly at the right pixel is an + independent sensor observation, not a report +Source C Sentinel-2 via CDSE: the first clear pass after the date, false-colour + B12/B8A/B4. A new burn scar or structural change on imagery acquired after + and absent before is the third, strongest leg +Baseline the same FIRMS box for the preceding year, to show the pixel does not + routinely register; the preceding Sentinel-2 pass, to show the scar is new +Negative GDELT timelinevol for the facility name, to see whether coverage preceded + the imagery — coverage before the satellite pass means the reporting is not + derived from the imagery and is genuinely separate +Record per leg: dataset and pinned version, exact query, retrieval timestamp, + SHA-256 of any downloaded raster, and the scene ID with its acquisition time +``` + +Two legs from the same family is one leg: ACLED and Liveuamap both code from media reporting, +GDELT and a news aggregator both index the same wires, MODIS and VIIRS are different sensors but +can be on the same platform pass. A sensor observation plus a reported event is a genuine pair, and +imagery plus a sensor observation is better. + +State the negative result as explicitly as the positive one. "No VIIRS detection in a 2 km box on +either date, and the first clear Sentinel-2 pass is nine days later" is a useful finding about the +limits of what can be established, and publishing it is what distinguishes this work from +advocacy. + +## Environment + +| Source | What it gives you | +| --- | --- | +| [Global Forest Watch](https://www.globalforestwatch.org/) | Deforestation alerts and tree-cover change, near-real-time, with an API behind it. | +| [NASA FIRMS](https://firms.modaps.eosdis.nasa.gov/) | Active fire and thermal anomaly detections, updated within three hours of overpass. | +| [Global Fishing Watch](https://globalfishingwatch.org/map) | Apparent fishing effort and vessel-behaviour events inferred from AIS. See [transport tracking](/sheets/osint/transport-tracking) for its API. | +| [Copernicus Browser](https://browser.dataspace.copernicus.eu/) | Free Sentinel-1/2/3 imagery with band combinations for burn scars, water, vegetation and earthworks. | +| [UNOSAT](https://unosat.org/products) | UN satellite analyses of humanitarian emergencies, already interpreted and published with methodology. | +| [Resource Watch](https://resourcewatch.org/) | 300+ environmental and human-wellbeing datasets, several of them real-time. | + ## Tool reference | Tool | What it does | Cost | @@ -115,6 +521,82 @@ Water / flooding: B8A, B11, B4 - **Documenting harm carries duty of care.** Graphic material needs handling policies, and identifying victims or witnesses can endanger them. Minimise what you publish. +## Worked example + +One datum: a coordinate and a date. `46.482, 30.724` on 11 March 2026, from a post claiming a +strike on a grain terminal. The coordinate is real and the returns below are illustrative; what +matters is which query answers which part of the claim. + +```bash +# 1. was anything coded as an event there +TOKEN=$(curl -s -X POST 'https://acleddata.com/oauth/token' \ + -d 'username=you@example.com' --data-urlencode 'password=PW' \ + -d 'grant_type=password' -d 'client_id=acled' -d 'scope=authenticated' | jq -r .access_token) + +curl -s -H "Authorization: Bearer $TOKEN" -G 'https://acleddata.com/api/acled/read' \ + --data-urlencode 'country=Ukraine' --data-urlencode 'admin1=Odesa' \ + --data-urlencode 'event_date=2026-03-10|2026-03-13' \ + --data-urlencode 'event_date_where=BETWEEN' \ + --data-urlencode 'fields=event_date|sub_event_type|location|latitude|longitude|geo_precision|fatalities|source' \ + | jq '.data[]' +# one event, 2026-03-11, sub_event_type "Shelling/artillery/missile attack", +# geo_precision 2, coordinates the city centroid rather than the terminal +``` + +Geo-precision 2 means ACLED located this to a town, not a point — so it corroborates that something +happened in Odesa that day, and says nothing about this coordinate. + +```bash +# 2. an independent sensor: thermal detections in a tight box around the terminal +curl -s "https://firms.modaps.eosdis.nasa.gov/api/area/csv/$MAP_KEY/VIIRS_SNPP_NRT/30.70,46.46,30.75,46.50/2/2026-03-11" +# two detections, 2026-03-11 ~22:40 UTC, frp 4.8 and 2.1, confidence n, daynight N +``` + +A night-time detection with that radiative power inside a 4 km box is a real heat source, and +nothing about it came from a news report. + +```bash +# 3. the baseline — does this pixel burn routinely +for d in 2025-03-01 2025-06-01 2025-09-01 2025-12-01 2026-01-01 2026-02-01; do + curl -s "https://firms.modaps.eosdis.nasa.gov/api/area/csv/$MAP_KEY/VIIRS_SNPP_SP/30.70,46.46,30.75,46.50/10/$d" | tail -n +2 + sleep 2 +done | wc -l +# 0 — no detections in that box across sixty sampled days of the preceding year +``` + +That is what converts the hotspot from an observation into an anomaly. + +```bash +# 4. imagery: find the first clear pass after the date +TOK=$(curl -s -X POST \ + 'https://identity.dataspace.copernicus.eu/auth/realms/CDSE/protocol/openid-connect/token' \ + -H 'content-type: application/x-www-form-urlencoded' \ + -d 'grant_type=client_credentials' -d 'client_id=ID' \ + --data-urlencode 'client_secret=SECRET' | jq -r .access_token) + +curl -s -X POST 'https://sh.dataspace.copernicus.eu/api/v1/catalog/1.0.0/search' \ + -H "Authorization: Bearer $TOK" -H 'Content-Type: application/json' \ + -d '{"collections":["sentinel-2-l2a"],"bbox":[30.70,46.46,30.75,46.50], + "datetime":"2026-03-08T00:00:00Z/2026-03-20T00:00:00Z","limit":20, + "fields":{"include":["id","properties.datetime","properties.eo:cloud_cover"]}}' \ + | jq '.features[] | {id, datetime:.properties.datetime, cloud:.properties["eo:cloud_cover"]}' +# 2026-03-09 cloud 3%, 2026-03-14 cloud 11% — a usable before/after pair +``` + +Pull both as B12/B8A/B4 GeoTIFFs through `/api/v1/process`, hash each, and record the scene IDs. +The 14 March image shows a dark scar and a collapsed roof span on one silo where the 9 March image +shows an intact structure. + +What is now defensible: a thermal anomaly at the coordinate on the night of 11 March, from a sensor +and not a report; no comparable detection in that pixel across the preceding year; and structural +change visible between two named Sentinel-2 scenes bracketing the date. Three legs, two of them +independent of all media reporting, each tied to a query, a retrieval time and a hash. + +What is not: the cause. Neither the hotspot nor the scar distinguishes a missile strike from a +drone, an accident or deliberate arson, and attribution needs the remnant. If a photograph of +debris surfaces, that goes to [OSMP](https://osmp.ngo/) for a markings comparison, and the +identification is stated as consistent-with until a lot number is legible. + ## Broader catalogues - [Conflict OSINT](https://tools.osintnewsletter.com/tool-categories/conflict-osint) diff --git a/src/content/sheets/osint/image-video-forensics.md b/src/content/sheets/osint/image-video-forensics.md @@ -4,9 +4,9 @@ description: "Read a file's metadata, test it for manipulation, and work out wha category: osint subcategory: "Images & Video" tags: [osint, forensics, metadata, exif, verification] -tools: [exiftool, ffprobe, fotoforensics, forensically] +tools: [exiftool, ffprobe, ffmpeg, mediainfo, fotoforensics, forensically, invid] difficulty: intermediate -updated: 2026-09-28 +updated: 2026-10-03 references: - name: "Bellingcat's Online Investigation Toolkit" url: "https://bellingcat.gitbook.io/toolkit" @@ -25,76 +25,325 @@ references: ## What this covers What the file itself says, as distinct from what its caption says. Metadata gives you camera, time -and sometimes location; encoder artefacts tell you about the processing history; manipulation -analysis suggests where to look harder. None of it is proof on its own, and all of it is easy to -over-read. +and sometimes location; container and encoder artefacts tell you the processing history; +manipulation analysis suggests where to look harder. None of it is proof on its own, and all of it +is easy to over-read. + +## Method + +1. **Hash the file before you touch it.** `shasum -a 256 evidence.jpg`, recorded with the time you + received it and who from. Every later claim you make is about that hash. Work on a copy. +2. **Ask for the original.** Every major platform re-encodes and strips metadata on upload, so a + file pulled from a timeline has already lost most of what you want. The original off the camera + or phone is a different artefact from the one in the post. +3. **Read metadata before anything else**, because it is cheap and it either corroborates the claim + or contradicts it. A contradiction is a lead; agreement is weak evidence, since metadata is + trivially written. +4. **Inventory the container separately from the codec.** The file extension, the container brand + and the codec inside it are three independent facts, and a mismatch between them means a tool + touched the file after capture. +5. **Only then run manipulation analysis.** Error Level Analysis and noise residual tell you where + to look, not what happened. Decide in advance what result would change your conclusion; + otherwise you will find whatever you went looking for. +6. **Corroborate outside the file.** The defensible finding is the one that survives a second, + independent source — a different upload of the same footage, a shadow angle consistent with the + claimed time, a satellite pass over the claimed place. The file alone almost never settles it. +7. **Record what you ran.** Keep the exact command, its full output and the tool version alongside + the hash. "exiftool showed no GPS" is not reproducible; a saved JSON dump with a version string + is. + +## Key tools + +### ExifTool + +The reference metadata reader and writer, and the only one worth learning properly. It parses more +tag families than anything else — EXIF, XMP, IPTC, ICC, maker notes, QuickTime atoms, PDF and +RIFF — and it tells you which family each tag came from, which matters because the same logical +field appears in several places and they disagree when a file has been edited. -## Metadata - -`exiftool` is the tool. Everything else is a wrapper. +```bash +brew install exiftool # or: apt install libimage-exiftool-perl +``` ```bash -# everything, grouped by tag family +# everything, grouped and labelled by the tag family it came from exiftool -a -G1 -s image.jpg -# the fields that usually matter -exiftool -Make -Model -DateTimeOriginal -CreateDate -GPSLatitude -GPSLongitude \ - -Software -Orientation image.jpg +# the fields that usually decide a question, in one line +exiftool -Make -Model -DateTimeOriginal -CreateDate -ModifyDate \ + -GPSLatitude -GPSLongitude -Software -Orientation image.jpg -# GPS as a decimal pair you can paste into a map +# every timestamp in the file at once: camera, GPS, filesystem, XMP +exiftool -time:all -a -G1 -s image.jpg + +# GPS as a decimal pair you can paste straight into a map exiftool -n -p '$GPSLatitude,$GPSLongitude' image.jpg -# recurse a directory into CSV for comparison across a set -exiftool -r -csv -Make -Model -DateTimeOriginal -GPSPosition ./images/ > meta.csv +# maker notes and undecoded tags — where lens, shutter count and burst IDs hide +exiftool -U -a -G1 image.jpg + +# structural sanity check: does the file match its own format spec +exiftool -validate -warning -a image.jpg + +# the full set as JSON, which is what you archive next to the hash +exiftool -j -g1 -a image.jpg > image.exif.json -# strip metadata before publishing (protect your own sources) +# a whole directory into one CSV, so you can sort a set by camera and time +exiftool -r -csv -Make -Model -SerialNumber -DateTimeOriginal -GPSPosition ./images/ > meta.csv + +# embedded metadata in video streams, not just the container header +exiftool -ee -G3 -a clip.mp4 + +# strip everything before you publish, to protect a source exiftool -all= -overwrite_original copy.jpg ``` -Reading it: +Reading the output: `Make`/`Model` should match the claimed device, and a `SerialNumber` that +recurs across a set ties those files to one body. `Software` naming an editor means the file was +processed — not that it was faked, but that this is not the original. `DateTimeOriginal` is capture +time from the camera clock, which carries no timezone and is routinely wrong by hours or years; +`CreateDate` and `ModifyDate` can be later. `-validate` flagging a truncated or non-conforming +structure is a genuine signal that something rewrote the file. + +Missing metadata is the normal case, not a red flag — platform stripping accounts for almost all of +it. The reverse error is worse: metadata that supports a convenient claim proves very little, +because `exiftool` writes as easily as it reads, and a forged `DateTimeOriginal` takes one command. +Treat supporting metadata as consistent-with, and contradicting metadata as a thread to pull. -- **`Make` / `Model`** should match the claimed device. A "phone photo" carrying a DSLR model is a - contradiction worth chasing. -- **`Software`** naming an editor means the file was processed. That is not manipulation, but it - means what you have is not the original. -- **Timestamps** come from the camera clock, which is often wrong and usually has no timezone. - `DateTimeOriginal` is capture; `CreateDate` and `ModifyDate` may be later. -- **GPS** is genuinely useful when present. It is also the first thing platforms strip. -- **Missing metadata is the normal case**, not a red flag. Every major platform strips EXIF on - upload. Absence tells you it went through a platform, nothing more. +### ffprobe -## Video +Stream and container inventory for video and audio. Use it to establish what the file actually is +before you argue about what it shows: how many streams, which codec, what container brand, whether +the frame rate is constant, and where the keyframes fall. ```bash -# full stream and container metadata -ffprobe -v quiet -print_format json -show_format -show_streams video.mp4 +brew install ffmpeg # ffprobe ships with ffmpeg +``` -# creation time and encoder, which often identify the app that produced the file -ffprobe -v quiet -show_entries format_tags=creation_time,encoder -of default=nw=1 video.mp4 +```bash +# full container and stream dump, archived next to the hash +ffprobe -v error -print_format json -show_format -show_streams video.mp4 > video.probe.json -# frame-level detail: is the frame rate constant, are there splice points -ffprobe -select_streams v -show_frames -show_entries frame=pkt_pts_time,pict_type \ - -of csv video.mp4 | head -50 +# the container's own identity: brand, duration, and who wrote it +ffprobe -v error -show_entries format=format_name,format_long_name,duration:format_tags \ + -of default=nw=1 video.mp4 + +# codec vs container in one view — mismatches start here +ffprobe -v error -select_streams v -show_entries \ + stream=codec_name,codec_tag_string,profile,level,pix_fmt,width,height,r_frame_rate,avg_frame_rate,nb_frames \ + -of default=nw=1 video.mp4 + +# encoder string and creation time, which usually name the producing app +ffprobe -v error -show_entries format_tags=creation_time,encoder,com.apple.quicktime.software \ + -of default=nw=1 video.mp4 + +# frame types and timestamps: is the cadence constant, where are the I-frames +ffprobe -v error -select_streams v -show_entries frame=pts_time,pict_type \ + -of csv=p=0 video.mp4 | head -50 + +# keyframe interval from the packet layer, cheaper than decoding frames +ffprobe -v error -select_streams v -show_entries packet=pts_time,flags,size \ + -of csv=p=0 video.mp4 | awk -F, '$2 ~ /K/ {print $1}' | head -20 + +# is there a second audio track, a timecode track, or stray metadata streams +ffprobe -v error -show_entries stream=index,codec_type,codec_name -of csv=p=0 video.mp4 ``` -Encoder strings are a strong fingerprint of the producing application. A file claiming to be -straight off a phone but carrying a desktop editor's encoder has been through an edit. +Note that `pkt_pts_time` was removed in ffmpeg 5 — on any current build the entry is `pts_time`, +and an old recipe using the former silently returns an empty column rather than erroring. -## Manipulation analysis +**Container-versus-codec mismatch is the single most useful manipulation signal in video.** A file +named `.mp4` whose `format_name` is `matroska`, an iPhone clip whose major brand is not `qt`/`mp4`, +H.264 in a container the capturing device never writes, or a `creation_time` later than the claimed +upload — each means something re-muxed or re-encoded the file. Encoder strings are a strong +fingerprint of the producing application: a clip claimed to be straight off a phone but carrying a +desktop editor's encoder has been through an edit, and that is a finding you can state plainly. +Variable frame rate where the device records constant, or an `nb_frames` that does not match +duration times frame rate, points the same way. + +What this does not tell you: nothing in the stream inventory distinguishes an innocent re-encode — +a messaging app, a download wrapper, a format conversion for upload — from a deliberate edit. It +tells you the file is not the original, and that you should be arguing about provenance rather than +about pixels. + +### ffmpeg + +Extraction, not analysis. Use it to pull the frames you will actually look at, and to cut the +segment you will cite, without silently re-encoding what you cite. + +```bash +# every keyframe, decoded fast because P- and B-frames are skipped entirely +ffmpeg -skip_frame nokey -i video.mp4 -fps_mode vfr keyframes/kf_%04d.png + +# frames at scene changes: 0.3-0.5 is the useful threshold range +ffmpeg -i video.mp4 -vf "select='gt(scene,0.4)',showinfo" -fps_mode vfr scenes/sc_%04d.png + +# a contact sheet of keyframes, for finding the shot you need +ffmpeg -skip_frame nokey -i video.mp4 -vf "scale=240:-1,tile=5x4" -frames:v 1 sheet.jpg + +# the exact frame at a timestamp, for a geolocation attempt +ffmpeg -ss 00:01:37.500 -i video.mp4 -frames:v 1 -q:v 2 frame.jpg + +# cut a clip without re-encoding, so the excerpt keeps the original stream +ffmpeg -ss 00:01:30 -to 00:01:50 -i video.mp4 -c copy excerpt.mp4 + +# extract the audio untouched — background sound often dates or places a clip +ffmpeg -i video.mp4 -vn -c:a copy audio.m4a + +# log scene-change scores to a file instead of writing frames, to find the cut points +ffmpeg -i video.mp4 -vf "select='gt(scene,0.3)',metadata=print:file=scenes.txt" -f null - +``` + +`-c copy` matters for anything you will publish or hand to someone else: re-encoding an excerpt +destroys the compression history that any later analysis would use, and makes your excerpt +unfalsifiable in the wrong direction. `-fps_mode vfr` replaced `-vsync vfr` in ffmpeg 5; both are +accepted on current builds but only the former is documented. + +Scene-change detection finds cuts, and cuts in material claimed to be a single continuous recording +are worth explaining. It also fires on pans, flashes and camera shake, so read the list as +candidates. + +### MediaInfo + +A second opinion on container structure, with a tag vocabulary closer to how broadcast and +camera manufacturers name things. Worth running alongside `ffprobe` because it surfaces writing +library, muxing application and per-track UUIDs that ffmpeg's inventory does not break out. + +```bash +brew install mediainfo # or: apt install mediainfo +``` + +```bash +# human-readable, every track +mediainfo video.mp4 + +# full detail including container atoms and writing library +mediainfo --Full video.mp4 + +# machine-readable, for archiving next to the hash +mediainfo --Output=JSON video.mp4 > video.mediainfo.json + +# just the fields that identify the producing tool +mediainfo --Inform="General;%Format%|%Encoded_Application%|%Encoded_Library%|%Encoded_Date%" video.mp4 + +# batch a directory and diff the writing libraries across a set +mediainfo --Output=CSV ./clips/ > clips.csv +``` + +Where `ffprobe` and `mediainfo` disagree about the container, that disagreement is itself +informative: it usually means the file is malformed or was written by something that does not +follow the spec. + +### FotoForensics + +Web-only Error Level Analysis, JPEG quality estimation and metadata, run on an upload. The useful +part is not the ELA image, it is the combination of the quality estimate and the metadata panel on +the same screen. + +What you do: paste a URL or upload the file at [fotoforensics.com](https://fotoforensics.com/). +Read the panels in this order. + +```text +1. Metadata confirm it matches what exiftool told you; if it differs, the + upload was re-encoded in transit and the analysis is of the wrong file +2. JPEG Quality a quality estimate well below the camera's normal output means + at least one re-save; "last saved at ~75%" is a provenance fact +3. ELA uniform noise across a single-source photo is the expected result; + look for a region whose edge brightness differs sharply from + identically-textured regions elsewhere in the same frame +4. Hidden Pixels cropped-out image data still present in the file +``` + +Capture for the record: the analysis URL, a screenshot of each panel, and the quality estimate as a +number. The URL alone is not enough, because the site's analysis can change. + +**This site publicly retains uploads.** Anything you submit becomes part of a public corpus. Do not +put a source's photograph, an unpublished image or anything identifying a person at risk into it — +use Forensically instead, which runs locally in the browser. + +**ELA is the most over-interpreted technique in this field.** Differing compression across regions +has mundane causes: resizing, text overlay, a logo burned in, successive saves, a screenshot of a +screenshot. ELA tells you where to look harder. It never tells you a region was pasted in, and an +ELA image with arrows drawn on it is not a finding. + +### Forensically + +The same family of analyses, running entirely client-side in your browser, so the file never leaves +your machine. That makes it the right default for sensitive material, and it adds clone detection +and a noise view that FotoForensics does not. + +Open [29a.ch/photo-forensics](https://29a.ch/photo-forensics/) and drop the file in. The panels +worth your time: + +```text +Clone Detection finds regions similar to other regions in the same image; raise the + threshold until repetitive texture (foliage, gravel, sky) stops matching + itself, then look at what still matches +Noise Analysis a pasted region often carries a different sensor noise floor; flat or + smoothed patches in an otherwise noisy frame are the signal +Level Sweep sweeps the brightness range to expose banding and edges hidden in + shadow or blown highlights +Luminance Gradient surfaces inconsistent lighting direction across objects that should + share one light source +Magnifier pixel-level inspection for resampling edges around inserted text +``` + +Shut the browser tab when you are done and nothing has been transmitted. Noise analysis degrades +badly on anything heavily compressed or resized, which is most images sourced from social media — +if the file has been through a platform, expect all of these panels to be inconclusive and say so +rather than straining to read them. + +### InVID / WeVerify plugin + +A browser extension that bundles the video-specific steps: keyframe extraction from a platform URL +without downloading the file yourself, reverse image search of those keyframes across several +engines in one click, a Twitter/X timeline search by time window, and a magnifier. + +Install from [weverify.eu/verification-plugin](https://weverify.eu/verification-plugin/). It is +maintained as part of an EU research project, so expect individual platform integrations to break +when those platforms change their markup; the keyframe and reverse-search functions are the durable +parts. Use it for triage, then come back to `ffmpeg` and `exiftool` for anything you intend to +publish, because the plugin does not give you a reproducible command or a hash. + +### Hashing and the record + +The thing that makes a forensic finding defensible is not the analysis, it is the chain from the +file you were given to the file you analysed. + +```bash +# the hash you cite, computed before any processing +shasum -a 256 evidence.mp4 | tee evidence.sha256 + +# hash a whole set, so a later re-download can be compared +find ./evidence -type f -exec shasum -a 256 {} \; > manifest.sha256 + +# verify nothing drifted while you worked +shasum -a 256 -c manifest.sha256 + +# perceptual hash, for matching re-encodes of the same footage across platforms +ffmpeg -i a.mp4 -i b.mp4 -filter_complex "signature=nb_inputs=2:detectmode=full" -f null - +``` + +Record alongside the hash: where the file came from and when, the archive copy +([see archiving and evidence](/sheets/osint/archiving-and-evidence)), the tool versions, and the +exact commands with their output. A cryptographic hash and a perceptual hash answer different +questions — the first proves you analysed the file you were given, the second lets you say two +differently-encoded uploads are the same recording. + +### Further viewers and one-pass analysers + +The long tail, useful when you want a second opinion or have no shell to hand. | Tool | What it does | | --- | --- | -| [FotoForensics](https://fotoforensics.com/) | Error Level Analysis, JPEG quality estimation, metadata. Note that it publicly retains uploads — do not use it on sensitive material. | -| [Forensically](https://29a.ch/photo-forensics/) | Clone detection, noise analysis, level sweep, magnifier. Runs client-side in the browser, so nothing is uploaded. | -| [Reveal Image Verification Assistant](https://mever.iti.gr/forensics/) | Multiple filters in one pass, with a report. | -| [Jimpl](https://jimpl.com/) | Quick browser-based EXIF and GPS viewer. | -| [xIFr](https://xifr.eu/) | Detailed EXIF viewer including maker notes. | - -**ELA is widely over-interpreted.** Differing compression across regions has many innocent causes — -resizing, text overlay, successive saves. Treat ELA as "look here", never as "this is fake". +| [Reveal Image Verification Assistant](https://mever.iti.gr/forensics/) | Runs several forensic filters in one pass and produces a report, which saves time on triage; the filters are the same family as Forensically's, so treat agreement between them as one result, not two. | +| [Jimpl](https://jimpl.com/) | Browser EXIF and GPS viewer with a map pin. Quick, and fine for a file you do not mind uploading. | +| [xIFr](https://xifr.eu/) | Detailed EXIF viewer that decodes maker notes, which is the one thing a casual viewer usually drops. | +| [metadata2go](https://www.metadata2go.com/) | Metadata for both images and video in a browser, when you cannot install anything. | +| [DeepFake-O-Meter](https://zinc.cse.buffalo.edu/ubmdfl/deep-o-meter/) | Academic ensemble of synthetic-media detectors. Reports a score, not a verdict, and false positives on compressed footage are common — never cite it alone. | -For anything sensitive, prefer the client-side tools. Uploading a source's photo to a public -analysis site that retains submissions can expose them. +Any of these that works by upload has the same exposure problem as FotoForensics: assume the file +is retained. For sensitive material the only safe tools are the local ones. ## Facial recognition @@ -141,6 +390,59 @@ one from nothing, and never treat a match as confirmation on its own. - **Facial recognition on found images** raises real risk of misidentifying someone. Corroborate before acting, and consider whether identification is necessary at all. +## Worked example + +A 40-second clip arrives by Telegram, captioned as filmed that morning on a phone at a named +street corner. + +```bash +# 1. hash first, work on a copy +shasum -a 256 clip.mp4 +# 9f2c...e41b — recorded with "received 2026-03-12 09:14 UTC from <source>" + +# 2. metadata: no EXIF, no GPS, one QuickTime atom +exiftool -a -G1 -s -ee clip.mp4 | grep -iE 'software|encoder|create|model' +# [QuickTime] CreateDate : 2026:03:11 22:41:08 +# [QuickTime] Encoder : Lavf60.16.100 +``` + +`Lavf` is ffmpeg's muxer, not a phone. The clip was written by a desktop or server-side tool, so +whatever it shows, this file is not the camera original — and `CreateDate` is the evening *before* +the claimed morning. + +```bash +# 3. container inventory: does the shape match a phone capture +ffprobe -v error -select_streams v -show_entries \ + stream=codec_name,codec_tag_string,pix_fmt,r_frame_rate,avg_frame_rate,nb_frames \ + -of default=nw=1 clip.mp4 +# codec_name=h264 codec_tag_string=avc1 pix_fmt=yuv420p +# r_frame_rate=30/1 avg_frame_rate=2997/100 nb_frames=1198 +``` + +Constant 30 declared, 29.97 actual — a broadcast-standard rate a phone does not record. Two +independent signals now say re-encoded. + +```bash +# 4. find the cuts, because the caption says one continuous take +ffmpeg -i clip.mp4 -vf "select='gt(scene,0.4)',metadata=print:file=scenes.txt" -f null - +# 3 scene changes at 11.2s, 24.6s, 31.0s + +# 5. pull the frame with the readable shopfront +ffmpeg -ss 00:00:26 -i clip.mp4 -frames:v 1 -q:v 2 frame26.jpg +``` + +Three cuts in a clip described as one take is the finding; the frame is the pivot. Reverse image +search on `frame26.jpg` ([see reverse image search](/sheets/osint/reverse-image-search)) returns +the same shopfront in a news photo dated fourteen months earlier, and +[geolocation](/sheets/osint/geolocation) puts the corner two streets from the one named. Shadow +direction in the frame is consistent with late afternoon, not morning. + +What you can defensibly say: the file was produced by ffmpeg, contains at least three edits, is at +least fourteen months old in part, and shows a different corner than claimed — each tied to a named +command, its output, and the hash `9f2c...e41b`. What you cannot say from the file alone is who +assembled it or why. That needs the second independent source: the earlier news photo, which is +what actually carries the date. + ## Broader catalogues - [Image and Video Analysis OSINT](https://tools.osintnewsletter.com/tool-categories/image-and-video-analysis-osint) diff --git a/src/content/sheets/osint/maps-and-satellite-imagery.md b/src/content/sheets/osint/maps-and-satellite-imagery.md @@ -4,7 +4,7 @@ description: "Choose the right imagery source for the question, find historical category: osint subcategory: "Geospatial" tags: [osint, satellite, maps, gis, street-view] -tools: [qgis, google-earth-pro, sentinel-hub, openstreetmap] +tools: [qgis, gdal, overpass, copernicus-data-space, google-earth-pro, mapillary, earthexplorer] difficulty: intermediate updated: 2026-09-28 references: @@ -28,6 +28,23 @@ Which imagery to use, how to get at older coverage, and the street-level sources The main skill is matching the source to the question — resolution, revisit frequency and archive depth trade off against each other and no single provider wins on all three. +## Method + +1. **Fix the question before the source.** "Has this compound grown since 2019" and "what is the + writing on that sign" are answered by different satellites, and no source answers both. +2. **Bound the area and the dates.** A coordinate with no radius and a year with no window will + produce thousands of scenes and no answer. Write both down first. +3. **Start free and coarse.** Sentinel-2 at 10m settles most change-detection questions. Spend + high-resolution effort only on confirming the specific thing it points at. +4. **Record the capture date of every frame you use**, from the source's own metadata rather than + the page you found it on. An undated image supports no time-sensitive claim. +5. **Cross-check the base layer.** Two providers disagree on building footprints, place names and + borders. Say which one you used, and look at a second before concluding a feature is new. +6. **Measure in a metric CRS.** Distances taken in degrees are wrong by a factor that varies with + latitude, and the tool will not warn you. +7. **Pivot to street level last.** Once satellite imagery has given you candidates, Mapillary or + Street View either confirms the ground detail or kills the candidate in seconds. + ## Choosing a source | Question | Source | @@ -50,7 +67,7 @@ accessible archive of high-resolution coverage. Note the imagery date shown at t the single most important piece of context and the most commonly ignored. **Landsat** goes back to 1972 and is the only free option for multi-decade change. Browse it via -[EarthExplorer](https://earthexplorer.usgs.gov/) or Sentinel Hub's Playground. +[EarthExplorer](https://earthexplorer.usgs.gov/), covered below. ## Street-level beyond Google @@ -63,40 +80,302 @@ Google's coverage is deep but not universal, and its capture dates are sometimes Always check several; a street with no Google coverage often has Mapillary imagery. -## OpenStreetMap as a database +## Key tools + +### Overpass API + +OpenStreetMap's query interface, and the reason OSM is a database rather than a picture. Give it a +tag combination and a bounding box and it returns the features — which is how you turn "a church +with a red roof next to a roundabout, somewhere in this province" into a list of candidates. +[Overpass Turbo](https://overpass-turbo.eu/) is the browser front end; the API behind it takes +`curl`, and automating it is what makes a large search tractable. + +```bash +# the public endpoint rejects requests without a User-Agent — this is the usual first failure +UA='osint-research/1.0' +OVERPASS='https://overpass-api.de/api/interpreter' + +# every pharmacy in a bounding box (south,west,north,east) +curl -s -A "$UA" "$OVERPASS" --data-urlencode \ + 'data=[out:json][timeout:25];node["amenity"="pharmacy"](52.37,4.88,52.38,4.90);out body;' + +# bridges over waterways — matching a described scene to candidate locations +curl -s -A "$UA" "$OVERPASS" --data-urlencode \ + 'data=[out:json][timeout:60];way["bridge"="yes"](52.3,4.8,52.4,4.95);out geom;' \ + | jq '.elements | length' + +# two features near each other: a mosque within 200m of a school +curl -s -A "$UA" "$OVERPASS" --data-urlencode \ + 'data=[out:json][timeout:90]; + way["amenity"="place_of_worship"]["religion"="muslim"](35.6,51.3,35.8,51.5)->.w; + node(around.w:200)["amenity"="school"];out center;' + +# cell masts, which are mapped far more completely than people expect +curl -s -A "$UA" "$OVERPASS" --data-urlencode \ + 'data=[out:json];node["man_made"="mast"]["tower:type"="communication"](48.1,16.3,48.3,16.5);out center;' + +# count first, then fetch: a national-scale query that returns nothing costs you nothing +curl -s -A "$UA" "$OVERPASS" --data-urlencode \ + 'data=[out:csv(::count)][timeout:120];area["ISO3166-1"="NL"]->.a;node["amenity"="fuel"](area.a);out count;' + +# straight to a file your GIS can open +curl -s -A "$UA" "$OVERPASS" --data-urlencode \ + 'data=[out:json][timeout:60];node["aeroway"="aerodrome"](34.0,43.0,35.0,44.5);out center;' \ + -o aerodromes.json +``` + +Queries are metered by server CPU time, not request count: a careless `[timeout:900]` over a whole +country will get you a 429 and then a temporary ban from the main instance. Run the count form +first. Because OSM is crowd-sourced, a feature's absence means nobody mapped it — never that it +is not there — and tag completeness varies by an order of magnitude between countries. For +proximity searches without writing Overpass QL by hand, Bellingcat's +[OpenStreetMap Search](https://osm-search.bellingcat.com/) wraps the same data. + +### GDAL + +The library every GIS tool is built on, and faster than any of them from a shell. Install +`gdal` with your package manager (`brew install gdal`, `apt install gdal-bin`). In OSINT work it +does four things: tell you what a file actually contains, crop and reproject it, georeference a +photograph or scanned map, and build band maths like NDVI without opening a GUI. + +```bash +# what is in this file: CRS, extent, pixel size, bands, and the capture metadata +gdalinfo -stats scene.tif | head -40 + +# crop to a bounding box in the file's own CRS, so you stop moving 2GB around +gdal_translate -projwin 4.88 52.38 4.90 52.37 -projwin_srs EPSG:4326 scene.tif crop.tif -OSM is queryable, which makes it far more powerful than its rendered map suggests. Overpass -Turbo takes queries like: +# reproject to Web Mercator for overlay on a slippy basemap +gdalwarp -t_srs EPSG:3857 -r cubic crop.tif crop_3857.tif + +# georeference a scanned map: four ground control points (pixel x, pixel y, lon, lat), then warp +gdal_translate -of GTiff -a_srs EPSG:4326 \ + -gcp 120 95 4.8855 52.3805 -gcp 1890 110 4.9015 52.3799 \ + -gcp 1875 1410 4.9010 52.3702 -gcp 135 1395 4.8860 52.3708 \ + scan.png scan_gcp.tif +gdalwarp -r cubic -t_srs EPSG:4326 -overwrite scan_gcp.tif scan_geo.tif + +# mosaic a directory of tiles into one virtual raster — no copying, instant +gdalbuildvrt mosaic.vrt tiles/*.tif + +# NDVI from Sentinel-2 bands: vegetation loss, burn scars, new earthworks +gdal_calc.py -A B08.jp2 -B B04.jp2 --outfile=ndvi.tif \ + --calc="(A.astype(float)-B)/(A.astype(float)+B+0.0001)" + +# hillshade from a DEM, for reading terrain in a photograph's background +gdaldem hillshade -z 2 dem.tif hillshade.tif + +# a shareable PNG at a sane size, with the world file so it stays georeferenced +gdal_translate -of PNG -outsize 25% 25% -co WORLDFILE=YES crop.tif preview.png +``` + +`gdalinfo` is the honesty check: if it reports no CRS, the file is a picture and any measurement +you take off it is invented. Georeferencing error concentrates away from your control points, so +put them at the corners of the area you care about and expect metres of error, not centimetres. +Resampling with `-r cubic` makes imagery look better and makes pixel-level forensics worse — use +`-r near` when the pixels themselves are the evidence. + +### QGIS + +The desktop GIS, and where a question stops being "look at this" and becomes "measure this, +against these layers". Free from [qgis.org](https://www.qgis.org). Add satellite imagery as a +basemap with **Layer → Add Layer → Add XYZ Layer** and one of these URLs: ```text -[out:json][timeout:25]; -// every pharmacy within the bounding box -node["amenity"="pharmacy"]({bbox}); -out body geom; +https://server.arcgisonline.com/ArcGIS/rest/services/World_Imagery/MapServer/tile/{z}/{y}/{x} +https://tile.openstreetmap.org/{z}/{x}/{y}.png +https://mt1.google.com/vt/lyrs=s&x={x}&y={y}&z={z} ``` +Note the `{z}/{y}/{x}` order on the Esri service and `{z}/{x}/{y}` on the others — swapped axes +are the reason a basemap loads as noise. The Georeferencer (**Layer → Georeferencer**) does the +same job as the `gdal_translate -gcp` run above with a point-and-click interface and a visible +residual error per point, which is worth the GUI on its own. + +Everything in Processing also runs headless, which is how a one-off analysis becomes repeatable: + +```bash +# every algorithm available, including the ones your installed plugins add +qgis_process list + +# the parameters for one algorithm, before you guess at them +qgis_process help native:buffer + +# a 200m buffer around candidate points +qgis_process run native:buffer -- INPUT=candidates.gpkg DISTANCE=200 OUTPUT=buffered.gpkg + +# reproject a layer to a metric CRS so distances mean something +qgis_process run native:reprojectlayer -- \ + INPUT=candidates.geojson TARGET_CRS='EPSG:3857' OUTPUT=candidates_3857.gpkg + +# centroids of building polygons, for matching against a geotagged photo set +qgis_process run native:centroids -- INPUT=buildings.gpkg OUTPUT=centroids.gpkg + +# keep only the features inside an area of interest +qgis_process run native:extractbylocation -- \ + INPUT=centroids.gpkg PREDICATE=0 INTERSECT=aoi.gpkg OUTPUT=inside.gpkg +``` + +Measurements in a geographic CRS (`EPSG:4326`) are in degrees, not metres, and QGIS will happily +give you a meaningless number. Reproject to a local metric CRS or a UTM zone before you measure +anything you intend to publish, and say which CRS you used. + +### Copernicus Data Space + +The free Sentinel archive, and the only no-cost source with a revisit frequency short enough for +change detection. Sentinel-2 gives 10m optical every five days; Sentinel-1 is radar and sees +through cloud. The catalogue is searchable without an account; downloading needs a free +registration. Note that Sentinel Hub Playground and the old EO Browser are retired — the browser +is now [Copernicus Browser](https://browser.dataspace.copernicus.eu/). + +```bash +CAT='https://catalogue.dataspace.copernicus.eu/odata/v1/Products' + +# what Sentinel-2 exists for a date window — no token needed for search +curl -s -G "$CAT" \ + --data-urlencode "\$filter=Collection/Name eq 'SENTINEL-2' and ContentDate/Start gt 2026-09-01T00:00:00.000Z and ContentDate/Start lt 2026-09-05T00:00:00.000Z" \ + --data-urlencode '$top=5' | jq -r '.value[] | "\(.ContentDate.Start) \(.Name)"' + +# the same, bounded to an area of interest +curl -s -G "$CAT" \ + --data-urlencode "\$filter=Collection/Name eq 'SENTINEL-2' and OData.CSC.Intersects(area=geography'SRID=4326;POLYGON((4.85 52.36,4.95 52.36,4.95 52.40,4.85 52.40,4.85 52.36))') and ContentDate/Start gt 2026-08-01T00:00:00.000Z" \ + --data-urlencode '$top=10' | jq -r '.value[].Name' + +# radar instead, for a cloudy week +curl -s -G "$CAT" \ + --data-urlencode "\$filter=Collection/Name eq 'SENTINEL-1' and ContentDate/Start gt 2026-09-01T00:00:00.000Z" \ + --data-urlencode '$top=5' | jq -r '.value[].Name' + +# a download needs a token from the Keycloak identity service +export ACCESS_TOKEN=$(curl -s -d 'client_id=cdse-public' -d "username=$CDSE_USER" \ + -d "password=$CDSE_PASS" -d 'grant_type=password' \ + 'https://identity.dataspace.copernicus.eu/auth/realms/CDSE/protocol/openid-connect/token' \ + | jq -r .access_token) + +# then fetch the product by its Id +curl -s -L -H "Authorization: Bearer $ACCESS_TOKEN" \ + "$CAT(08f7cbba-56c6-4730-b2d1-63ee8a5b9536)/\$value" -o product.zip +``` + +Filter on `Collection/Name` or the query will be slow enough to time out. A Sentinel-2 product +name encodes the tile and the processing level — `L2A` is atmospherically corrected and what you +want for comparing two dates; `L1C` is not. The cloud percentage in the metadata is for the whole +100km tile, so a "12% cloud" scene can still be solid cloud over your target. Tokens expire in +minutes; re-request rather than caching them. + +### Google Earth Pro + +Web and desktop, free, and still the most accessible archive of sub-metre historical imagery. The +desktop build is the one worth having, because the browser version hides the two features that +matter. + +The historical slider (**View → Historical Imagery**, or the clock icon) steps through every +coverage date for the view. Work it deliberately: note the date stamp in the status bar for every +frame you use, because that stamp is the whole evidentiary value of the screenshot. The elevation +profile (draw a path, then **Edit → Show Elevation Profile**) answers "could you see X from Y", +which is a common verification question that imagery alone cannot settle. + +```text +Ruler tool line and path lengths, plus a bearing in degrees +Polygon + area rooftop and compound areas, for matching a described structure +Show Elevation Profile line-of-sight and terrain between two points +Sun / lighting slider shadow direction at a chosen date and time +Import KML/GPX drop in Overpass output or a track for overlay +Status-bar date stamp the capture date of the imagery currently drawn +``` + +The displayed date is the date of the *dominant* image in the view; a mosaic can blend captures +months apart, with a visible seam. Zooming changes which image is drawn, so the date can change +under you without the view appearing to move. 3D buildings are models, not imagery, and are not +evidence of anything. + +### Mapillary + +Crowd-sourced street-level imagery, frequently covering roads Google never drove and often more +recent. Free API, and unlike Street View it is queryable by bounding box and capture date, which +makes "what did this junction look like in March" a scriptable question. + +```bash +# token from mapillary.com/dashboard/developers — the prefix is literally "MLY|" +TOKEN='MLY|xxxx|xxxx' +API='https://graph.mapillary.com' + +# images in a bounding box (minLon,minLat,maxLon,maxLat) — must be under 0.01 degrees square +curl -s -H "Authorization: OAuth $TOKEN" \ + "$API/images?bbox=4.890,52.372,4.895,52.375&fields=id,captured_at,compass_angle,geometry&limit=50" \ + | jq -r '.data[] | "\(.captured_at) \(.id)"' + +# only images from a date window, which is the point of using this over Street View +curl -s -H "Authorization: OAuth $TOKEN" \ + "$API/images?bbox=4.890,52.372,4.895,52.375&start_captured_at=2026-03-01T00:00:00Z&end_captured_at=2026-04-01T00:00:00Z&fields=id,captured_at" + +# one image in full, including the direction the camera faced and the original file +curl -s -H "Authorization: OAuth $TOKEN" \ + "$API/IMAGE_ID?fields=id,captured_at,compass_angle,camera_type,thumb_2048_url,geometry" + +# the detected objects in a frame — traffic signs are the usable ones for geolocation +curl -s -H "Authorization: OAuth $TOKEN" \ + "$API/images?bbox=4.890,52.372,4.895,52.375&fields=id,detections.value" + +# download the frame you are going to cite +curl -s -H "Authorization: OAuth $TOKEN" "$API/IMAGE_ID?fields=thumb_2048_url" \ + | jq -r .thumb_2048_url | xargs curl -s -o frame.jpg +``` + +The bounding box limit is real: anything larger than 0.01 degrees square is rejected rather than +truncated, so tile your area. `compass_angle` is the camera bearing and is what lets you say which +side of the street a feature is on. Coverage is contributor-driven, so a road can have 2019 and +2026 imagery and nothing between, and sequence positions are GPS traces — metres of error in +cities, more in canyons. + +### EarthExplorer + +Web only, free with registration, and the only practical route to two archives nothing else +carries: Landsat back to 1972, and declassified US reconnaissance imagery from the 1960s to the +1980s at resolutions that still surprise people. + +Search at [earthexplorer.usgs.gov](https://earthexplorer.usgs.gov/). The interface is four tabs +and the order matters: + ```text -[out:json][timeout:25]; -// bridges over waterways in the area — useful for matching a described scene -way["bridge"="yes"]({bbox}); -out geom; +1 Search Criteria draw a polygon or enter coordinates, then set the date range +2 Data Sets Landsat → Landsat Collection 2 Level-2 for analysis-ready scenes + Declassified Data → Declass 1/2/3 for CORONA, ARGON, KH-7 and KH-9 +3 Additional Criteria cloud cover ceiling, and the sensor for declassified missions +4 Results the footprint icon shows coverage; the browse icon previews it ``` -For relationship searches without writing Overpass by hand, use Bellingcat's OpenStreetMap Search -or Spot, which wrap the same data in a proximity-search interface. +Register before you search, because the download buttons are hidden until you log in, and read +the scene's entity ID — it encodes the mission and date, and it is what you cite. Declassified +frames arrive as scanned film with no georeferencing at all, which is where the +`gdal_translate -gcp` workflow above earns its keep. Landsat Collection 2 Level-2 is corrected and +comparable between dates; Level-1 is not, so do not difference the two. -## QGIS +### SunCalc and ShadeMap -The free desktop GIS, and what you want once a question involves more than looking. Typical OSINT -uses: loading satellite basemaps as XYZ tiles, georeferencing a photograph or a scanned map against -known points, measuring distances and bearings, and building a map for publication. +Web only, free, and the fastest way to put a time on a photograph you have already geolocated. A +shadow's direction and length at a known place is a clock, and these two read it in opposite +directions. -Add an XYZ basemap with Layer → Add Layer → Add XYZ Layer, for example Esri World Imagery: +[SunCalc](https://www.suncalc.org/) takes a coordinate and a date and draws the sun's azimuth and +elevation through the day; you match the shadow bearing in the image to the time that produces it. +[ShadeMap](https://shademap.app) does the inverse, simulating the shadows that buildings and +terrain actually cast at a chosen moment, which is what you need in a city where the shadow comes +off a tower rather than the subject. ```text -https://server.arcgisonline.com/ArcGIS/rest/services/World_Imagery/MapServer/tile/{z}/{y}/{x} +Input: coordinate, date, and the shadow bearing measured off the image +Read: the time or times of day whose azimuth matches that bearing +Check: shadow LENGTH against solar elevation — bearing alone gives two candidate times +Capture: the coordinate, date, computed azimuth and elevation, and the tool's own URL +Caveat: a date you have not independently established makes the whole result circular ``` +Both assume you have the location right; a 50m error in position barely moves the azimuth, but a +wrong date moves it by degrees per week near the solstices. Shadow work narrows a time, it does +not prove one — pair it with [image and video forensics](/sheets/osint/image-video-forensics) and +with whatever the metadata claims. + ## Tool reference | Tool | What it does | Cost | @@ -171,6 +450,38 @@ https://server.arcgisonline.com/ArcGIS/rest/services/World_Imagery/MapServer/til - **Basemap labels disagree**, particularly on disputed borders and place names. Say which source you used. +## Worked example + +One datum: the coordinate **52.3738, 4.8909**, pulled from a photograph's caption, and a claim +that a warehouse there was demolished in the summer of 2026. + +1. **Establish what is mapped.** An Overpass query for `building` ways in a small box around the + coordinate returns three polygons, one tagged `building=warehouse` with an `addr:street`. That + gives the feature a name and an address to search on. +2. **Find free imagery either side of the claim.** The Copernicus OData catalogue, filtered to + `SENTINEL-2` and intersected with a polygon around the point, lists L2A products for late May + and early September 2026. Two dates bracketing the claim is the minimum useful set. +3. **Crop and compare.** `gdal_translate -projwin` cuts both scenes to the same 500m box, + `gdalwarp -t_srs EPSG:3857` puts them in the same CRS, and the pair opened in QGIS shows the + roof present in May and bare ground in September. At 10m the roof is six pixels across — enough + to see it go, not enough to see how. +4. **Confirm at resolution.** Google Earth Pro's historical slider over the same point has a + July 2026 frame at sub-metre scale showing partial demolition and plant on site. Record the + status-bar date, not the date you looked. +5. **Check the ground.** Mapillary `images?bbox=…&start_captured_at=2026-08-01T00:00:00Z` returns + a contributor sequence from August with `compass_angle` facing the plot: hoarding up, structure + gone. Street level dates the end of the work more precisely than any satellite pass. +6. **Measure, then say so.** Reprojected to UTM, the QGIS measure tool puts the cleared footprint + at 1,840 m², consistent with the OSM polygon's area. Quote the CRS alongside the number. +7. **Time the photograph, if it matters.** The caption claims mid-July. SunCalc for the coordinate + on 14 July returns an azimuth matching the shadow bearing in the image at around 17:40 local — + consistent, not proof, and recorded as such. + +What you can assert: a structure present on a dated 10m scene in May and absent in September, with +a sub-metre frame in July showing demolition in progress and street-level imagery in August +showing it finished. What you cannot: who did it, or why — no imagery source on this page +carries that. + ## Broader catalogues - [Geolocation and Maps OSINT](https://tools.osintnewsletter.com/tool-categories/geolocation-and-maps-osint) diff --git a/src/content/sheets/osint/people-search.md b/src/content/sheets/osint/people-search.md @@ -4,7 +4,7 @@ description: "Registries, court records, aggregators and breach data for identif category: osint subcategory: "People & Identity" tags: [osint, people, public-records, breach-data] -tools: [pipl, intelx, dehashed, hibp] +tools: [hibp, intelx, aleph, courtlistener, opensanctions, ratsit] difficulty: intermediate updated: 2026-09-28 references: @@ -53,8 +53,299 @@ unusually productive: - **Norway / Finland / Denmark** — equivalents exist; tax records are partly public in Norway. - **Nigeria** — NigeriaPhonebook for telephone-to-name lookups. - **United States** — county-level court and property records are the real source; national - aggregators are convenience layers over them. See Search Systems for a directory of the - underlying databases. + aggregators are convenience layers over them. + [Search Systems](https://www.searchsystems.net/) is a directory of the underlying databases, + organised by state and record type. + +## Key tools + +### Have I Been Pwned + +The reference for which breach corpora hold an address, and the only tool in this area with a +stable documented API. Account lookups sit behind a paid key now — the cheapest tier is a few +dollars a month — while the breach catalogue and the password range endpoint stay free and need +no key at all. + +```bash +# the breach catalogue: free, no key, useful for dating a corpus +curl -s 'https://haveibeenpwned.com/api/v3/breaches' | jq -r '.[].Name' | head + +# one breach in detail — when it happened, how big, what fields leaked +curl -s 'https://haveibeenpwned.com/api/v3/breach/Adobe' | jq '{BreachDate,PwnCount,DataClasses}' + +# which breaches hold this address; both headers are mandatory +curl -s 'https://haveibeenpwned.com/api/v3/breachedAccount/target@example.com?truncateResponse=false' \ + -H "hibp-api-key: $HIBP_KEY" -H 'user-agent: osint-research' \ + | jq -r '.[] | "\(.BreachDate) \(.Name)"' + +# paste sites that carried the address, which often predate the breach being named +curl -s 'https://haveibeenpwned.com/api/v3/pasteAccount/target@example.com' \ + -H "hibp-api-key: $HIBP_KEY" -H 'user-agent: osint-research' + +# every breached address under a domain you have proven you control +curl -s 'https://haveibeenpwned.com/api/v3/breachedDomain/example.com' \ + -H "hibp-api-key: $HIBP_KEY" -H 'user-agent: osint-research' + +# stealer-log hits — infostealer output, so far more recent than breach corpora (Pro tier) +curl -s 'https://haveibeenpwned.com/api/v3/stealerLogsByEmail/target@example.com' \ + -H "hibp-api-key: $HIBP_KEY" -H 'user-agent: osint-research' + +# check a password without sending it: only the first five SHA-1 characters leave your machine +printf 'hunter2' | shasum | tr 'a-f' 'A-F' | cut -c1-5 \ + | xargs -I{} curl -s "https://api.pwnedpasswords.com/range/{}" | head -3 +``` + +Omitting `user-agent` returns 403 rather than a useful error. A 429 carries a `retry-after` +header in seconds; honour it, because the limit is per key and hammering it gets the key +suspended. The output tells you a corpus containing that address was published — not that your +subject created the account, not that they still use the address, and not what the password was. + +### Intelligence X + +Web-first, and the broadest selector search available without a corporate contract: it indexes +leaks, darknet pages, document dumps and its own historical web crawl, and it searches by +selector type rather than free text. Reach for it when an identifier returns nothing elsewhere. + +Paste the selector at [intelx.io](https://intelx.io/) and read the result pane by bucket — the +left-hand facets split hits into leaks, darknet, web and documents, and the bucket matters more +than the hit count. Selectors it understands: + +```text +target@example.com email address +example.com domain, including subdomain hits +198.51.100.0/24 CIDR range ++14155550123 phone number, E.164 +1BoatSLRHtKNngkdXEeobR76b53LETtpyT cryptocurrency address +d41d8cd98f00b204e9800998ecf8427e hash or document selector +``` + +For the record, capture the item's **date added** and its bucket, not just the snippet — +Intelligence X keeps items long after the source is gone, so the snippet alone cannot be re-verified. The API +takes the key from the developer tab of your account in an `x-key` header; phonebook (bulk +selector extraction) and most export volume are paid tiers, and the free tier is metered tightly +enough that you will feel it inside an hour. + +### OCCRP Aleph and Library of Leaks + +Aleph is a document-and-entity index built for cross-border corruption work: leaked archives, +scraped registries, sanctions lists and court filings normalised into one searchable entity +graph. It beats a people-search aggregator whenever the subject appears in documents rather than +directories. [Library of Leaks](https://search.libraryofleaks.org/) is the same software over a +different, leak-heavy corpus. + +```bash +pipx install alephclient + +# register free, then take the key from your Aleph user profile — anonymous API calls get a 401 +export ALEPHCLIENT_HOST=https://aleph.occrp.org +export ALEPHCLIENT_API_KEY=REPLACE_ME + +# people matching a name, across every dataset you can read +curl -s -G 'https://aleph.occrp.org/api/2/entities' \ + -H "Authorization: ApiKey $ALEPHCLIENT_API_KEY" \ + --data-urlencode 'q=Sergei Ivanov' --data-urlencode 'filter:schema=Person' \ + | jq -r '.results[] | "\(.properties.name[0]) \(.collection.label)"' + +# companies instead of people — the usual pivot off a director's name +curl -s -G 'https://aleph.occrp.org/api/2/entities' \ + -H "Authorization: ApiKey $ALEPHCLIENT_API_KEY" \ + --data-urlencode 'q=Sergei Ivanov' --data-urlencode 'filter:schema=Company' | jq '.total' + +# which datasets exist, so you can say what you did and did not search +curl -s 'https://aleph.occrp.org/api/2/collections?limit=50' \ + -H "Authorization: ApiKey $ALEPHCLIENT_API_KEY" \ + | jq -r '.results[] | "\(.foreign_id) \(.label)"' + +# one entity in full, including the documents it was extracted from +curl -s 'https://aleph.occrp.org/api/2/entities/ENTITY_ID' \ + -H "Authorization: ApiKey $ALEPHCLIENT_API_KEY" | jq '.properties' + +# push your own case documents in so they are indexed alongside the public corpora +alephclient crawldir --foreign-id my-case-2026 ./case-documents +``` + +Names in Aleph are transliterated inconsistently because the sources are, so run the variants +from [username and account work](/sheets/osint/usernames-and-accounts) before concluding a name +is absent. An entity is an extraction from a document, which means both the spelling and the +role can be wrong; open the source document before you cite it. Uploading case documents sends +them to someone else's server — check that is acceptable for your material first. + +### CourtListener and RECAP + +Free, keyless-to-browse access to US federal dockets, opinions and PACER filings that other +researchers have paid for and donated. For a US subject this is the best single documentary +source on this page, because a filing names addresses, employers, counsel and relatives in one +place. + +```bash +# token from your free account profile; note the literal word "Token" +export CL_TOKEN=REPLACE_ME + +# federal dockets mentioning a name (type=r gives dockets with nested documents) +curl -s -G 'https://www.courtlistener.com/api/rest/v4/search/' \ + -H "Authorization: Token $CL_TOKEN" \ + --data-urlencode 'q="Dana Whitfield"' --data-urlencode 'type=r' \ + | jq -r '.results[] | "\(.dateFiled) \(.court) \(.caseName)"' + +# case law opinions instead of dockets +curl -s -G 'https://www.courtlistener.com/api/rest/v4/search/' \ + -H "Authorization: Token $CL_TOKEN" \ + --data-urlencode 'q="Dana Whitfield"' --data-urlencode 'type=o' | jq '.count' + +# narrow to one court and a date window +curl -s -G 'https://www.courtlistener.com/api/rest/v4/search/' \ + -H "Authorization: Token $CL_TOKEN" \ + --data-urlencode 'q=Whitfield' --data-urlencode 'type=d' \ + --data-urlencode 'court=txwd' --data-urlencode 'filed_after=2020-01-01' + +# judges, for checking who heard a case +curl -s -G 'https://www.courtlistener.com/api/rest/v4/search/' \ + -H "Authorization: Token $CL_TOKEN" \ + --data-urlencode 'q=Whitfield' --data-urlencode 'type=p' + +# the docket itself, including the parties +curl -s 'https://www.courtlistener.com/api/rest/v4/dockets/?id=12345' \ + -H "Authorization: Token $CL_TOKEN" | jq '.results[0] | {case_name,date_filed,assigned_to_str}' + +# page with the cursor the previous response handed back, not an offset +curl -s -G 'https://www.courtlistener.com/api/rest/v4/search/' \ + -H "Authorization: Token $CL_TOKEN" \ + --data-urlencode 'q=Whitfield' --data-urlencode 'type=r' --data-urlencode 'cursor=CURSOR' +``` + +RECAP holds what somebody has already fetched from PACER, so absence means nobody bought that +document, not that the case does not exist — check PACER's own index before saying there is no +case. Search results are cached for ten minutes, so polling for new filings wastes your quota; +use the Alert API for monitoring. State courts are mostly absent, which is where most US +litigation actually happens. + +### OpenSanctions + +Sanctions lists, politically-exposed-person data and criminal watchlists, merged and deduplicated +into one entity model with a matching API. Use it to answer "is this person on a list" properly, +rather than eyeballing a dozen national PDFs — and to record a clean negative. + +```bash +export OS_KEY=REPLACE_ME + +# full-text search, for when you only have a name +curl -s -G 'https://api.opensanctions.org/search/default' \ + -H "Authorization: ApiKey $OS_KEY" --data-urlencode 'q=Sergei Ivanov' \ + | jq -r '.results[] | "\(.caption) \(.schema) \(.datasets|join(","))"' + +# proper screening: describe the entity and let the matcher score candidates +curl -s -X POST 'https://api.opensanctions.org/match/default' \ + -H "Authorization: ApiKey $OS_KEY" -H 'Content-Type: application/json' \ + -d '{"queries":{"q1":{"schema":"Person","properties":{"name":["Sergei Ivanov"],"birthDate":["1965"],"country":["ru"]}}}}' \ + | jq -r '.responses.q1.results[] | "\(.score) \(.caption) \(.datasets|join(","))"' + +# one entity in full, with every list that carries it and why +curl -s 'https://api.opensanctions.org/entities/NK-A7z1Z' \ + -H "Authorization: ApiKey $OS_KEY" | jq '{caption,datasets,properties}' + +# which lists are in the index, so your negative result has a defined scope +curl -s 'https://api.opensanctions.org/catalog' | jq -r '.datasets[].name' | head -30 +``` + +A date of birth or a country code moves the score far more than extra name tokens; a bare name +query on a common name returns confident nonsense. The score is a similarity, not a finding — +anything you publish needs the underlying list entry read by a human. Trial keys come free with +a business email and public-interest keys are granted for journalism and research; if the metering +gets in the way, the matching engine (`yente`) is open source and self-hostable against the same +free bulk data. + +### EDGAR full-text search + +Free, keyless, and routinely forgotten in people work: every SEC filing since 2001 is full-text +searchable, and filings name directors, officers and beneficial owners with addresses. A +Form D or a proxy statement is a dated, signed document naming a person — which is exactly the +kind of source an aggregator hit is not. + +```bash +# the SEC requires a descriptive User-Agent with contact details; without one you get blocked +UA='osint-research research@example.com' + +# every filing mentioning a name +curl -s -A "$UA" 'https://efts.sec.gov/LATEST/search-index?q=%22Dana+Whitfield%22' \ + | jq -r '.hits.hits[] | "\(._source.file_date) \(._source.display_names[0])"' + +# restrict to a form type — Form D names the directors of a private placement +curl -s -A "$UA" 'https://efts.sec.gov/LATEST/search-index?q=%22Dana+Whitfield%22&forms=D' + +# date-bounded, which is how you tie a person to a company in a given year +curl -s -A "$UA" \ + 'https://efts.sec.gov/LATEST/search-index?q=%22Dana+Whitfield%22&dateRange=custom&startdt=2019-01-01&enddt=2021-12-31' + +# every filing by one company, from its CIK +curl -s -A "$UA" 'https://data.sec.gov/submissions/CIK0000320193.json' \ + | jq '{name,addresses,formerNames}' +``` + +The browsable version is at [sec.gov/edgar/search](https://www.sec.gov/edgar/search/), and it is +the right place to read a hit. Full-text coverage starts in 2001; older filings are indexed by +company but not searchable by content. A name in a filing is the name as filed, so middle +initials and suffixes are inconsistent across years. + +### Ratsit + +Web only, Sweden only, and far more revealing than anything equivalent elsewhere: registered +address, year and month of birth, income bracket, property and vehicle flags, and who else is +registered at the address. Swedish population-register data is public, so this is a legitimate +primary source rather than a broker's guess. + +Search the name at [ratsit.se](https://www.ratsit.se/), narrow by county when the name is common, +and read the profile page. For the record, capture these fields and nothing inferred from them: + +```text +Folkbokförd adress registered address, i.e. where the state believes they live +Född birth year and month (day is withheld on the free view) +Inkomst / Taxeringsår income bracket AND the assessment year it belongs to +Fordon / Fastighet vehicle and property flags, each a pointer to another register +Personer på adressen others registered at the same address — household, not family +``` + +The income figure is always a year or two behind and is meaningless without the assessment year +printed beside it. + +**Search logged out.** Ratsit sells subjects a paid feature listing the logged-in users who +looked them up; anonymous searches leave no entry in that log. Mrkoll, the nearest competitor, +shows subjects a profile-view count but not identities. Either way, an account turns a passive +lookup into something the subject can notice. + +### US people-search aggregators + +[FastPeopleSearch](https://www.fastpeoplesearch.com/), [That's Them](https://thatsthem.com/) and +[Radaris](https://radaris.com/) are free, web only, and the fastest route to a candidate address +and phone for a US subject. They are also data brokers reselling scraped and purchased records of +unknown age. + +Enter the name plus a city or state — a bare name on a common surname is unusable. Read the +result page as three separate claims with separate reliability: current address, phone numbers, +and "possible relatives". The first two are checkable against a county property or voter record; +the third is inference from shared addresses and surnames and is wrong often enough that +repeating it is reckless. + +```text +Dana Whitfield Austin TX name + city: the only query form that works +Dana Whitfield 78701 name + ZIP, when the city name is ambiguous +(512) 555-0123 reverse phone, usually the highest-precision entry point +1200 Example St, Austin TX reverse address, to enumerate the household +``` + +Capture the page with an archiving tool, because broker pages change without notice — see +[archiving and evidence](/sheets/osint/archiving-and-evidence). + +These sites are also the mechanism by which your subject's data is public in the first place; most +carry an opt-out form, and a subject who has used it will be absent from the broker while still +present in the underlying county record. Absence here means nothing. + +### Pipl + +**Sales-gated since its pivot to enterprise identity resolution.** There is no self-serve tier, +no web search for individuals, and contracts start in the thousands of dollars a year. Treat the +Pipl entry in the table below as historical. For the same job without a contract, the live +options are the registry-first route above, Intelligence X for selector search, and OSINT +Industries or Osintly if you need a commercial aggregator with per-query pricing. ## Breach and leak data @@ -72,6 +363,13 @@ and a phone number. Handle carefully: this is other people's stolen data. Do not put a live target's credentials into a third-party lookup you do not control, and do not treat a password from a breach as authorisation to use it. +DeHashed is the one worth a note on access: it is credit-metered now, and the old +basic-auth `api.dehashed.com/search?query=` endpoint that circulates in every scripted wrapper is +retired. The current API is v2, documented behind a login at +[dehashed.com/docs](https://dehashed.com/docs), so check the endpoint there rather than copying a +wrapper. The web search works without any of +that. Leak-Lookup's free tier is a hash-only check; its searchable API is paid. + ## Tool reference | Tool | What it does | Cost | @@ -88,7 +386,7 @@ treat a password from a breach as authorisation to use it. | [Person Lookup](https://personlookup.co.za/) | find individuals, phonenumbers, and adresses | free | | [Pipl](http://pipl.com/) | Identity information for professionals | paid | | [Ratsit](https://www.ratsit.se/) | Look up phone numbers/names (Sweden) | free | -| [Search Systems](https://publicrecords.searchsystems.net/) | Finding public record information online in over 70,000 databases organized by type and location to help you find property, criminal, court, birth, death… | free | +| [Search Systems](https://www.searchsystems.net/) | Finding public record information online in over 70,000 databases organized by type and location to help you find property, criminal, court, birth, death… | free | | [Skopenow](https://www.skopenow.com/) | Social Media Investigations - name, phone, email, username searches. | paid | | [Spokeo](http://spokeo.com/) | People search through email, phone, name | paid | | [Swedish Name Register](https://scb.se/hitta-statistik/sverige-i-siffror/namnsok/) | Find out how common a name is in Sweden based on census data | free | @@ -111,6 +409,59 @@ treat a password from a breach as authorisation to use it. collection of personal data about EU residents engages the GDPR regardless of the data being public. +## Worked example + +One datum: the name **Dana Whitfield**, typed under a signature on a PDF that also carries a +Travis County, Texas address block. Nothing else. + +1. **Establish the jurisdiction's real sources.** [Search Systems](https://www.searchsystems.net/) + lists the Travis County clerk, district clerk and appraisal district portals. That tells you + where the authoritative property and civil records live before you touch an aggregator. +2. **Federal dockets.** `type=r` search on CourtListener for `"Dana Whitfield"` returns one 2021 + docket in the Western District of Texas. The docket's party block gives a street address and + the name of a law firm — two new pivots, from a dated court document. +3. **Filings.** EDGAR full-text search for the same name, `forms=D`, returns a 2022 Form D naming + her as a director of a private company, with a business address in Austin and a company CIK. +4. **Company side.** The CIK into `data.sec.gov/submissions/CIK….json` gives the company's former + names and address history, which is how you find out the entity was renamed in 2023 — the + reason a plain name search missed it. +5. **Screening, including the negative.** The OpenSanctions `/match/default` call with + `name`, `country=us` and the birth year inferred from the docket returns no result above 0.6. + Record that as a checked negative with the datasets listed, not as "nothing found". +6. **Identifier confirmation.** The business email in the Form D goes into HIBP + `breachedAccount`. Two corpora dated 2019 and 2021 hold it, which confirms the address was in + use across that window — and nothing more. +7. **Aggregator, last.** That's Them on "Dana Whitfield, Austin TX" returns three candidate + addresses, one matching the docket. The other two, and the whole "possible relatives" block, + stay in the notes as unverified inference. + +The whole run, as commands: + +```bash +UA='osint-research research@example.com' +export CL_TOKEN=REPLACE_ME OS_KEY=REPLACE_ME HIBP_KEY=REPLACE_ME + +curl -s -G 'https://www.courtlistener.com/api/rest/v4/search/' -H "Authorization: Token $CL_TOKEN" \ + --data-urlencode 'q="Dana Whitfield"' --data-urlencode 'type=r' | jq -r '.results[].caseName' + +curl -s -A "$UA" 'https://efts.sec.gov/LATEST/search-index?q=%22Dana+Whitfield%22&forms=D' \ + | jq -r '.hits.hits[] | "\(._source.file_date) \(._source.display_names[0])"' + +curl -s -A "$UA" 'https://data.sec.gov/submissions/CIK0001234567.json' | jq '{name,formerNames}' + +curl -s -X POST 'https://api.opensanctions.org/match/default' -H "Authorization: ApiKey $OS_KEY" \ + -H 'Content-Type: application/json' \ + -d '{"queries":{"q1":{"schema":"Person","properties":{"name":["Dana Whitfield"],"country":["us"]}}}}' \ + | jq -r '.responses.q1.results[] | "\(.score) \(.caption)"' + +curl -s 'https://haveibeenpwned.com/api/v3/breachedAccount/dana@example.com' \ + -H "hibp-api-key: $HIBP_KEY" -H 'user-agent: osint-research' | jq -r '.[].Name' +``` + +What you can assert: a named person, tied by two independent dated documents to one company and +one city, with one address corroborated by a court filing. What you cannot: anything the +aggregator alone said, and any relationship nobody filed. + ## Broader catalogues - [People OSINT](https://tools.osintnewsletter.com/tool-categories/people-osint) diff --git a/src/content/sheets/osint/social-media-platforms.md b/src/content/sheets/osint/social-media-platforms.md @@ -4,7 +4,7 @@ description: "What each major platform still exposes to an unauthenticated resea category: osint subcategory: "Social Media" tags: [osint, social-media, telegram, tiktok] -tools: [telepathy, yt-dlp, instaloader] +tools: [yt-dlp, gallery-dl, instaloader, telethon, arctic-shift, atproto] difficulty: intermediate updated: 2026-09-28 references: @@ -29,6 +29,26 @@ per-platform picture: what is still reachable without an account, what needs one survive contact with the current APIs. Assume anything here can break without notice — platforms change access rules faster than tooling keeps up. +## Method + +1. **Decide whether you can be seen before you touch anything.** Some platforms attribute a view, + a follow or a group join to your account. Work out which, on this platform, today — then pick + the account you are willing to burn. +2. **Archive first, analyse second.** Capture the post, the profile and the surrounding thread + before you start pulling threads. Accounts get deleted mid-investigation, routinely. +3. **Take the public surface without an account** where it exists: profile metadata, follower + counts, pinned posts, video metadata. Much of this is still reachable unauthenticated and + costs you nothing. +4. **Pivot on identifiers, not display names.** Numeric IDs, DIDs and channel IDs survive renames + and vanity-URL changes; handles and names do not. +5. **Trace to origin.** The account with the most engagement is almost never the source. Follow + forwards, quote-posts and reuploads backwards until the chain stops. +6. **Date everything from the platform, not the page.** Upload timestamps, message IDs and + creation dates are platform-side facts; the date shown on a reposted screenshot is not. +7. **Expect the tooling to be broken.** Anything built on an undocumented endpoint dies without + notice. Check the project's last commit before you trust its output, and cross-check one result + by hand. + ## Telegram The most productive platform for open-source research, because public channels are genuinely @@ -41,19 +61,8 @@ public, history is retained, and forwarding metadata exposes the network between - **Channel creation dates and message IDs** are sequential, which lets you estimate when a channel started and how much has been deleted. -[Telepathy](https://github.com/proseltd/Telepathy-Community) is the standard toolkit for archiving -channels, mapping forwards and exporting membership. - -```bash -pipx install telepathy - -telepathy -t channelname # basic channel scrape -telepathy -t channelname -c # comprehensive: members, forwards, media -telepathy -t channelname -f 2024-01-01 -``` - -It needs Telegram API credentials from [my.telegram.org](https://my.telegram.org/) tied to a real -number, so use a number you are willing to burn. +Archiving a channel, mapping its forwards and checking a number against the platform all have +tooling below under [Key tools](#key-tools). ## Facebook @@ -69,30 +78,15 @@ Graph search is gone and most enumeration is dead. What still works: - Profile pictures, bios and post counts are visible without login; the feed usually is not. - `?__a=1` style JSON endpoints have been closed off repeatedly — assume they do not work. -- [Instaloader](https://instaloader.github.io/) still handles public profile and hashtag downloads, - including comments and geotags where present. - -```bash -pipx install instaloader - -instaloader profile_name # public posts + metadata -instaloader --no-pictures --comments profile_name -instaloader "#hashtag" -``` - -Aggressive use gets the account and IP rate-limited quickly, then blocked. +- [Instaloader](https://instaloader.github.io/) still handles public profile and hashtag + downloads, including comments and geotags where present; it has its own section below. +- Aggressive use gets the account and the IP rate-limited quickly, then blocked outright. ## X / Twitter Heavily restricted since the API changes. Advanced search still works while logged in, and is -still the best tool on the platform: - -```text -from:username since:2024-01-01 until:2024-06-30 -"exact phrase" filter:images -filter:retweets -geocode:51.5074,-0.1278,5km -url:example.com -``` +still the best tool on the platform — the operator set is below under +[Key tools](#key-tools). Archived snapshots are now often more reliable than the live site for deleted content — go to the [Wayback Machine](https://web.archive.org/) first. @@ -106,20 +100,451 @@ Archived snapshots are now often more reliable than the live site for deleted co ## YouTube -Still the most researcher-friendly large platform: +Still the most researcher-friendly large platform: metadata, channel listings, comments and +auto-generated subtitles all come out of `yt-dlp`, covered below. + +Upload timestamps are in the metadata and are a reliable earliest-possible date for footage. + +## Reddit + +- Profiles, posts and comment trees are readable without an account, and appending `.json` to any + Reddit URL returns the same thing as structured data. +- Deleted and removed content is the point of interest, and Reddit does not serve it. The + historical archives do — Arctic Shift below is the live successor to Pushshift. +- Subreddit moderator lists, wiki pages and automoderator configs are public and frequently + forgotten by the people who wrote them. + +## Bluesky + +Structurally the most open platform here, because the AT Protocol is a public API by design: every +post, follow and profile is a record in a repository you can read without an account. + +- **DIDs are permanent.** A handle is a pointer; the `did:plc:…` identifier survives every rename, + and the PLC directory keeps an audit log of every change to it — including the account's first + operation, which dates its creation. +- **Follow graphs are fully enumerable**, which makes network analysis possible in a way it is not + anywhere else on this page. +- **Search is the one thing that is not open.** Post search now needs an authenticated session; + profile and feed reads do not. + +## Key tools + +### yt-dlp + +Does far more than download: it reads the platform's own metadata for a video, a channel or a +playlist, and it covers well over a thousand sites, so the same invocation works on YouTube, +TikTok, Twitch, Instagram and most of the long tail. For dating footage and enumerating a +channel's output it is the first tool to reach for. ```bash -# metadata only, no download -yt-dlp --dump-json 'https://youtube.com/watch?v=VIDEOID' +pipx install yt-dlp + +# everything the platform will tell you about one video, without downloading it +yt-dlp --dump-json 'https://youtube.com/watch?v=VIDEOID' | jq '{upload_date,uploader_id,duration,view_count}' -# subtitles, including auto-generated, for keyword searching -yt-dlp --write-auto-subs --sub-langs en --skip-download URL +# just the fields you need, printed — faster to read than the full JSON +yt-dlp -O '%(upload_date)s %(uploader_id)s %(title)s' 'https://youtube.com/watch?v=VIDEOID' -# full channel listing -yt-dlp --flat-playlist --dump-json 'https://youtube.com/@channel/videos' +# the whole channel as a listing, without touching the videos themselves +yt-dlp --flat-playlist --dump-json 'https://youtube.com/@channel/videos' \ + | jq -r '[.upload_date,.id,.title] | @tsv' + +# auto-generated subtitles, which turn hours of video into greppable text +yt-dlp --write-auto-subs --sub-langs en --skip-download 'https://youtube.com/watch?v=VIDEOID' + +# comments, saved alongside the metadata, for identifying participants +yt-dlp --write-info-json --write-comments --skip-download 'https://youtube.com/watch?v=VIDEOID' + +# the thumbnail set, which is a reverse-image-search input and sometimes an earlier frame +yt-dlp --write-thumbnail --skip-download 'https://youtube.com/watch?v=VIDEOID' + +# one clip out of a long stream, by timestamp, instead of the whole six hours +yt-dlp --download-sections '*01:12:30-01:14:00' 'https://youtube.com/watch?v=VIDEOID' + +# a TikTok video with its metadata — same tool, same flags +yt-dlp --write-info-json 'https://www.tiktok.com/@user/video/1234567890123456789' + +# items 1 to 50 of a playlist, with an archive file so a resumed run does not re-fetch +yt-dlp -I 1:50 --download-archive seen.txt 'https://youtube.com/playlist?list=PLID' + +# the logged-in surface, when a public one does not exist — uses your browser's session +yt-dlp --cookies-from-browser firefox --dump-json URL ``` -Upload timestamps are in the metadata and are a reliable earliest-possible date for footage. +`upload_date` is the platform's own value and is a reliable earliest-possible date for footage; +it is not the date the footage was shot. Fields vary by extractor, so check with `--dump-json` +before scripting against a key. Passing `--cookies-from-browser` attaches your real session to +every request, which attributes the activity to that account — use a research profile. Extractors +break when platforms change, so `pipx upgrade yt-dlp` before blaming the URL. + +### gallery-dl + +The image-and-gallery counterpart to yt-dlp: it walks a profile, a tag or a thread and takes the +media plus the per-item JSON. Where it earns its place is bulk profile capture with metadata +intact — the captions, timestamps and IDs that make a media set evidence rather than a folder of +pictures. + +```bash +pipx install gallery-dl + +# what it would fetch, and the metadata for each item, without downloading +gallery-dl -j 'https://www.instagram.com/username/' + +# the available metadata keys for a URL, so your filters reference real fields +gallery-dl -K 'https://www.reddit.com/user/username/submitted/' + +# media plus a JSON sidecar per file, into a named directory +gallery-dl --write-metadata -D ./capture/username 'https://twitter.com/username/media' + +# direct URLs only, to hand to an archiving tool instead of downloading yourself +gallery-dl -g 'https://www.reddit.com/r/subreddit/comments/abc123/' + +# the first 50 items of a long profile, which is usually all you need to characterise it +gallery-dl --range 1-50 'https://www.flickr.com/photos/username/' + +# only the large images, skipping avatars and thumbnails +gallery-dl --filter 'image_width >= 1000' 'https://example-gallery.test/user/x' + +# an archive file, so re-running a monitored profile fetches only what is new +gallery-dl --download-archive seen.sqlite3 'https://www.instagram.com/username/' + +# slow and polite: the alternative is a 429 and then a block +gallery-dl --sleep 3 --sleep-request 2 --sleep-429 120 'https://www.instagram.com/username/' + +# a logged-in session where the public surface is gone +gallery-dl --cookies-from-browser firefox 'https://www.instagram.com/username/' +``` + +Per-site options live in a config file (`gallery-dl --config-create` writes a starter), and most +rate-limit problems are solved there rather than on the command line. Timestamps in the sidecar +JSON are the platform's, which is the point — but the field name differs per extractor, so read +`-K` output before you build a timeline. A profile capture is a snapshot: items deleted before you +ran it are simply absent, and nothing in the output tells you that. + +### Instaloader + +Instagram specifically, and still working where most Instagram tooling has died. Public profiles, +hashtags, locations and stories, with geotags and comments where they exist. + +```bash +pipx install instaloader + +# a public profile: posts, captions and metadata JSON +instaloader profile_name + +# metadata only — no media, much faster, much less conspicuous +instaloader --no-pictures --no-videos profile_name + +# comments and geotags, which are the parts worth having +instaloader --comments --geotags profile_name + +# just the profile picture at full resolution, for reverse image search +instaloader --no-posts profile_name + +# posts within a date window, using a Python filter expression +instaloader --post-filter 'date_utc >= datetime(2026,1,1)' profile_name + +# a hashtag rather than an account +instaloader '#hashtag' + +# stories and highlights need a session; both attribute the view to that account +instaloader --login YOUR_RESEARCH_ACCOUNT --stories --highlights profile_name + +# resume a monitored profile from where you stopped +instaloader --fast-update --latest-stamps stamps.ini profile_name +``` + +**Viewing a story is attributed.** The account whose session you used appears in the poster's +viewer list, so `--stories` is never a passive operation. Unauthenticated use is rate-limited +aggressively and then blocked by IP; authenticated use gets the account flagged and sometimes +disabled. Geotags are only present where the poster added them, and the location name is +user-chosen, not GPS. + +### Telegram: Telepathy and Telethon + +[Telepathy](https://github.com/proseltd/Telepathy-Community) is the toolkit the Telegram sections +of every OSINT guide point at, and it still runs — but it has been **unmaintained since July +2024**, with 2.3.4 the last release and the maintainers saying so in the README. Use it knowing +that; for anything you need to keep working, the durable route is Telethon, which tracks the +Telegram API itself. + +```bash +pipx install telepathy + +telepathy -t channelname # basic channel scrape +telepathy -t channelname -c # comprehensive: messages, members, media +telepathy -t channelname -c -f # plus a forward edgelist, the useful part +telepathy -u username # look up a user +telepathy -e # export your account's chat list +``` + +Both need API credentials from [my.telegram.org](https://my.telegram.org/), which are tied to a +real phone number — use one you are willing to lose. Telethon in twenty lines does the part that +matters, and does not rot: + +```python +# pip install telethon +from telethon.sync import TelegramClient + +with TelegramClient('research-session', API_ID, API_HASH) as client: + entity = client.get_entity('channelname') + print(entity.id, entity.title, entity.date) # channel id and creation date + + # message history, newest first; message IDs are sequential, so gaps are deletions + for msg in client.iter_messages(entity, limit=500): + fwd = msg.forward.chat.username if msg.forward and msg.forward.chat else None + print(msg.id, msg.date.isoformat(), fwd or '-', (msg.message or '')[:80]) + + # the forward graph: which channels this one amplifies, with counts + from collections import Counter + sources = Counter( + m.forward.chat.username + for m in client.iter_messages(entity, limit=2000) + if m.forward and m.forward.chat + ) + print(sources.most_common(20)) + + # membership, where the group exposes it — you appear in their list too + for user in client.iter_participants(entity, limit=200): + print(user.id, user.username, user.first_name) +``` + +Sequential message IDs are the quiet win here: a gap between IDs is deleted content, and the first +ID dates the channel. Joining a group to read it adds your research account to a list other members +can see, and scraping at speed from one account is exactly the behaviour Telegram bans for. Rate +limits surface as `FloodWaitError` with a duration — sleep for it rather than retrying. + +### telegram-phone-number-checker + +Bellingcat's tool for the one Telegram question with a clean answer: is this phone number +registered, and if so what does the account show. Also takes usernames. + +```bash +pip install telegram-phone-number-checker + +# credentials go in a .env: API_ID, API_HASH, PHONE_NUMBER +telegram-phone-number-checker --phone-numbers +14155550123 + +# several at once, comma-separated +telegram-phone-number-checker --phone-numbers +14155550123,+14155550124 + +# pull the profile photos too, for reverse image search +telegram-phone-number-checker --phone-numbers +14155550123 --download-profile-photos + +# by username instead of number +telegram-phone-number-checker --usernames johndoe + +# both in one run, to a named output file +telegram-phone-number-checker --phone-numbers +14155550123 --usernames johndoe --output results.json + +# through a proxy, since the lookups come from your account +pip install 'telegram-phone-number-checker[proxy]' +telegram-phone-number-checker --phone-numbers +14155550123 --proxy socks5://127.0.0.1:1080 +``` + +Checking a number works by adding it to your account's contacts, which is a write operation on +Telegram's side: the README's own advice is not to use a personal account, and that is not +boilerplate. A negative result means the number is not registered *or* the account's privacy +settings hide it from non-contacts — the tool cannot tell you which. See +[email and phone OSINT](/sheets/osint/email-and-phone) for the rest of the number work. + +### tiktok-hashtag-analysis + +Bellingcat's bulk primitive for TikTok: collect the posts carrying one or more hashtags, store the +metadata, and chart which hashtags co-occur. That co-occurrence graph is how a coordinated +campaign becomes visible, which individual video analysis never shows. + +```bash +pip install tiktok-hashtag-analysis + +# collect posts for one or more hashtags +tiktok-hashtag-analysis london paris + +# hashtags from a file, when the list is long +tiktok-hashtag-analysis --file hashtags.txt --output-dir ./tiktok-run + +# the top 20 co-occurring hashtags, as a table +tiktok-hashtag-analysis london --number 20 --table + +# the same, plotted +tiktok-hashtag-analysis london --number 20 --plot + +# download the videos too, capped per hashtag +tiktok-hashtag-analysis london --download --limit 200 + +# watch the browser work, which is how you debug a run that returns nothing +tiktok-hashtag-analysis london --headed -v +``` + +It drives a real browser, so it is slow and it breaks when TikTok changes its front end — a run +that returns zero posts for a busy hashtag means the scraper is broken, not that the hashtag is +empty. Check with `--headed` before concluding anything. Hashtag collection is a sample, never the +complete set, so counts are comparative at best. For a single video's exact upload time, +[Bellingcat's TikTok timestamp extractor](https://bellingcat.github.io/tiktok-timestamp) decodes it +from the video ID, and `yt-dlp --dump-json` gets the rest. + +### Arctic Shift + +The live successor to Pushshift for Reddit: a public, keyless archive of posts and comments, +including content since deleted from Reddit itself. This is the only practical way to read a +removed comment or reconstruct a user's history after they wiped it. + +```bash +AS='https://arctic-shift.photon-reddit.com/api' + +# everything an author posted, oldest first +curl -s -G "$AS/posts/search" --data-urlencode 'author=some_user' \ + --data-urlencode 'sort=asc' --data-urlencode 'limit=100' \ + | jq -r '.data[] | "\(.created_utc) \(.subreddit) \(.title)"' + +# their comments, which are usually more revealing than their posts +curl -s -G "$AS/comments/search" --data-urlencode 'author=some_user' \ + --data-urlencode 'limit=100' | jq -r '.data[] | "\(.subreddit) \(.body[0:120])"' + +# keyword search inside a subreddit and a date window +curl -s -G "$AS/posts/search" --data-urlencode 'subreddit=worldnews' \ + --data-urlencode 'title=wuhan' --data-urlencode 'after=2019-12-30' --data-urlencode 'limit=10' + +# only the fields you want, which keeps a wide sweep manageable +curl -s -G "$AS/posts/search" --data-urlencode 'subreddit=osint' \ + --data-urlencode 'fields=title,created_utc,author' --data-urlencode 'limit=100' + +# posts linking a specific domain — how a site spreads across subreddits +curl -s -G "$AS/posts/search" --data-urlencode 'url=example-news-daily.com' \ + --data-urlencode 'limit=100' | jq -r '.data[].subreddit' | sort | uniq -c | sort -rn + +# the full comment tree under one post, up to 25,000 comments +curl -s -G "$AS/comments/tree" --data-urlencode 'link_id=abc123' + +# known IDs straight to records, up to 500 per call +curl -s -G "$AS/comments/ids" --data-urlencode 'ids=t1_aaa,t1_bbb' +``` + +Keyword search on `title`, `selftext` or `body` requires an accompanying `author`, `subreddit`, +`link_id` or `parent_id` — a bare keyword sweep is refused, and it is refused for very active +authors and subreddits too. The archive holds what it ingested at the time, so an edit after +ingestion is invisible and a post removed within seconds of being made may never have been +captured. A record here is evidence that text was published, not that it is still live, and the +live check is a separate step. + +### Bluesky and the AT Protocol + +No key, no account, and a genuinely public API — the most researcher-friendly surface of any +current platform. Everything below runs against the public AppView except post search, which now +requires an authenticated session. + +```bash +API='https://public.api.bsky.app/xrpc' + +# handle to DID: the identifier that survives renames +curl -s "$API/com.atproto.identity.resolveHandle?handle=bellingcat.com" | jq -r .did + +# the PLC audit log: every handle and server change, and the creation date in entry one +curl -s 'https://plc.directory/did:plc:z72i7hdynmk6r22z27h6tvur/log/audit' \ + | jq -r '.[] | "\(.createdAt) \(.operation.alsoKnownAs // [] | join(","))"' + +# profile: display name, counts, and the account's own description +curl -s "$API/app.bsky.actor.getProfile?actor=bellingcat.com" \ + | jq '{handle,displayName,followersCount,postsCount,createdAt}' + +# the author's posts, paged with the cursor the response returns +curl -s "$API/app.bsky.feed.getAuthorFeed?actor=bellingcat.com&limit=100" \ + | jq -r '.feed[] | "\(.post.indexedAt) \(.post.record.text[0:100])"' + +# only posts carrying media, which is the usual filter for verification work +curl -s "$API/app.bsky.feed.getAuthorFeed?actor=bellingcat.com&limit=50&filter=posts_with_media" + +# the follow graph, fully enumerable in both directions +curl -s "$API/app.bsky.graph.getFollows?actor=bellingcat.com&limit=100" | jq -r '.follows[].handle' +curl -s "$API/app.bsky.graph.getFollowers?actor=bellingcat.com&limit=100" | jq -r '.followers[].handle' + +# a whole thread, replies included, from one post's AT URI +curl -s "$API/app.bsky.feed.getPostThread?uri=at://did:plc:xxxx/app.bsky.feed.post/yyyy" + +# post search needs a session — this returns 403 on the public host +curl -s "$API/app.bsky.feed.searchPosts?q=osint&limit=5" +``` + +The PLC audit log is the find here: it is an append-only record of every handle the account has +used and every server it has moved between, dated, with the first entry establishing when the +account was created. `indexedAt` is when the relay saw the post, not when the author wrote it — +`record.createdAt` is the author's claim and is trivially spoofable, so quote both. For search, +authenticate with `com.atproto.server.createSession` against the account's own PDS and send the +returned access token as a bearer; that turns a keyless read into an attributable one. + +### Meta Ad Library API + +The one Meta surface that is genuinely open, and the only reliable way to get spend and reach data +for political and issue advertising. Searchable by keyword or page, with the creative, the dates, +the platforms and the demographic breakdown. + +```bash +# access needs a Meta developer app, ID verification and Ad Library API approval first +TOKEN="$META_TOKEN" +G='https://graph.facebook.com/v23.0/ads_archive' + +# keyword search; ad_reached_countries is mandatory, and so is search_terms or search_page_ids +curl -s -G "$G" -d "access_token=$TOKEN" \ + -d 'search_terms=climate' -d "ad_reached_countries=['GB']" \ + -d 'fields=page_name,ad_delivery_start_time,ad_snapshot_url' | jq -r '.data[].page_name' | sort -u + +# one page's entire ad history, which is the attribution-relevant query +curl -s -G "$G" -d "access_token=$TOKEN" \ + -d 'search_page_ids=123456789' -d "ad_reached_countries=['GB']" \ + -d 'ad_active_status=ALL' \ + -d 'fields=ad_creative_bodies,ad_delivery_start_time,ad_delivery_stop_time,spend,impressions' + +# who paid, which is the field the transparency rules exist for +curl -s -G "$G" -d "access_token=$TOKEN" \ + -d 'search_terms=election' -d "ad_reached_countries=['GB']" \ + -d 'fields=bylines,page_name,spend,currency,publisher_platforms' + +# where it ran and to whom +curl -s -G "$G" -d "access_token=$TOKEN" \ + -d 'search_terms=election' -d "ad_reached_countries=['GB']" \ + -d 'fields=delivery_by_region,demographic_distribution,languages' + +# a date window, for tying a campaign to an event +curl -s -G "$G" -d "access_token=$TOKEN" \ + -d 'search_terms=referendum' -d "ad_reached_countries=['IE']" \ + -d 'ad_delivery_date_min=2026-01-01' -d 'ad_delivery_date_max=2026-03-31' \ + -d 'fields=page_name,spend,ad_snapshot_url' + +# page the result set with the cursor Graph hands back +curl -s -G "$G" -d "access_token=$TOKEN" -d 'search_terms=climate' \ + -d "ad_reached_countries=['GB']" -d 'fields=page_name' -d 'limit=100' -d 'after=CURSOR' +``` + +Leaving out both `search_terms` and `search_page_ids` is a parameter error, not an unfiltered +dump. `search_terms` matches the ad's own language and is capped at 100 characters, and spaces are +treated as AND. `spend` and `impressions` come back as ranges, not numbers — report them as +ranges. Coverage is political and issue ads in the countries where Meta is required to disclose +them; commercial ads appear in the web Ad Library with far less metadata and are not in the API +the same way. + +### X / Twitter advanced search + +Web only in practice. The API is priced out of reach for research, `snscrape` has been +non-functional against X since the 2023 access changes, and the public Nitter instances are +largely dead — so the live tool is X's own advanced search, which still works while logged in. + +```text +from:username since:2026-01-01 until:2026-06-30 one account, bounded by date +to:username replies directed at an account +"exact phrase" -filter:retweets the original post rather than its echoes +filter:images filter:links posts carrying media or outbound links +url:example-news-daily.com who linked a domain +geocode:51.5074,-0.1278,5km posts placed near a coordinate +min_faves:500 min_retweets:100 the versions that actually travelled +list:12345678 "phrase" search inside a list you have built +conversation_id:1234567890123456789 one thread, including replies +``` + +Every one of these requires a logged-in session, which attributes the searching to that account — +use a research account, and assume the queries are logged. Date-bounded `from:` queries are the +most reliable form; free-text search silently drops older results. For deleted posts, go to the +[Wayback Machine](https://web.archive.org/) and archive.today first: for X specifically, archived +snapshots are now a more dependable source than the live site. ## Tool reference @@ -186,6 +611,41 @@ Upload timestamps are in the metadata and are a reliable earliest-possible date - **Deleted is not gone, and present is not original.** Archive first, then analyse. - **Reposts dominate.** The account with the most engagement is rarely the origin. Trace back. +## Worked example + +One datum: the handle **@civicwatchnow**, read off a screenshot of a post, with no platform +attached and no other context. + +1. **Find the platforms.** A username sweep — see + [username and account discovery](/sheets/osint/usernames-and-accounts) — returns hits on + Bluesky, YouTube, Reddit and a Telegram channel of the same name. Four surfaces, four different + amounts of openness. +2. **Date the accounts, keylessly.** `resolveHandle` gives the Bluesky DID; the PLC audit log's + first entry dates the account to March 2026 and shows one earlier handle. An account presenting + itself as an established watchdog, three months old, is the first real finding. +3. **Characterise the output.** `yt-dlp --flat-playlist --dump-json` on the YouTube channel lists + 41 videos, all uploaded within six weeks, with `upload_date` clustered on weekdays. Volume and + cadence, from the platform's own metadata. +4. **Read what was deleted.** Arctic Shift's `comments/search?author=civicwatchnow` returns 300 + comments across eleven subreddits, including fourteen that no longer exist on Reddit. The + removed ones carry the same link, which the live profile does not show. +5. **Map the amplification.** The Telethon forward counter over the Telegram channel's last 2,000 + messages returns a top-twenty list in which three channels account for most forwards, and the + channel's own `entity.date` matches the Bluesky creation month. +6. **Check for paid distribution.** The Meta Ad Library `search_terms=civicwatch` with + `ad_reached_countries` set returns two advertisers running the same creative, with `bylines` + naming a company and `spend` as a range. That is a funding lead the organic surfaces never gave. +7. **Trace to origin, then archive.** The earliest instance of the clip in the original screenshot + turns out to be the TikTok upload, dated from the video ID rather than the display text; the + YouTube and Telegram copies are later. Capture all four profiles and the ad snapshots before + writing anything — see [archiving and evidence](/sheets/osint/archiving-and-evidence). + +What you can assert: four accounts under one handle, all created within a month of each other, +pushing a single clip whose earliest copy is dated on TikTok, amplified by three Telegram channels +and promoted by two named advertisers. What you cannot: that one person runs all four. A shared +handle is a lead until content, timing or a paid-distribution record ties them — and here it is the +ad byline, not the handle, doing the work. + ## Broader catalogues - [Social Media OSINT](https://tools.osintnewsletter.com/tool-categories/social-media-osint) diff --git a/src/content/sheets/osint/transport-tracking.md b/src/content/sheets/osint/transport-tracking.md @@ -4,9 +4,9 @@ description: "Track flights and ships live and historically, and understand the category: osint subcategory: "Transport" tags: [osint, aviation, maritime, adsb, ais] -tools: [adsbexchange, flightradar24, marinetraffic, opensky] +tools: [opensky, pyopensky, adsb-lol, adsbexchange, readsb, flightradar24, marinetraffic, equasis, global-fishing-watch] difficulty: intermediate -updated: 2026-09-28 +updated: 2026-10-03 references: - name: "Bellingcat's Online Investigation Toolkit" url: "https://bellingcat.gitbook.io/toolkit" @@ -24,70 +24,391 @@ references: ## What this covers -Aircraft and vessels broadcast their own positions, which makes transport one of the richest OSINT -domains. The important knowledge is not which site to open — it is where the data has gaps, because -the gaps are where the interesting behaviour is. +Aircraft and vessels broadcast their own positions, which makes transport one of the few OSINT +domains with a structured, queryable, global feed behind it. The important knowledge is not which +website to open — it is where the data has gaps, because the gaps are where the interesting +behaviour is, and distinguishing a coverage gap from an evasion is the whole skill. + +## Method + +1. **Pin the durable identifier first.** For aircraft that is the ICAO 24-bit hex address; for + vessels the IMO number. Registrations, callsigns, names and flags all change, frequently on + purpose, and a track assembled on a callsign will silently merge two aircraft. +2. **Look at live data on a map to orient yourself**, then stop using the map. Anything you intend + to assert needs the underlying records, because a screenshot of a map is not reproducible and + the free sites retain only days. +3. **Pull the track from an API and store it.** Positions, timestamps and the receiver metadata if + you can get it. Historical ADS-B and AIS are the parts that cost money, so capture while you can + see it rather than planning to come back. +4. **Resolve the operator separately from the registered owner.** Aircraft sit in trusts and + leasing companies; ships sit under flags of convenience with a manager in a third country. The + tracking feed names neither — that is a + [corporate records](/sheets/osint/companies-and-finance) problem. +5. **Characterise the gap before you call it a gap.** Establish the coverage baseline for that area + and that hour, from a second receiver network if possible, then show the vessel or aircraft was + inside coverage and absent anyway. Without that comparison you have a missing data point, not a + dark period. +6. **Corroborate position with something that is not a transponder.** Satellite imagery over the + claimed location and time, a port-call record, a photograph with a spotter's timestamp. A + transponder reports what it is told to report. ## Aircraft -Aircraft transmit ADS-B: identity, position, altitude, speed. Receivers are crowd-sourced, so -coverage follows population. +Aircraft transmit ADS-B: identity, position, altitude, speed, on 1090 MHz. Reception is +crowd-sourced, so coverage follows volunteer receivers, which follow population and wealth. | Source | Why use it | | --- | --- | -| [ADS-B Exchange](https://globe.adsbexchange.com/) | **Does not honour blocking requests.** The only major aggregator that shows aircraft others hide — which is precisely why it matters for research. | +| [ADS-B Exchange](https://globe.adsbexchange.com/) | **Does not honour blocking requests.** The only major aggregator that shows aircraft others hide, which is precisely why it matters for research. The map is free; the API is not. | | [Flightradar24](https://www.flightradar24.com/) | Best coverage and UI; filters blocked aircraft. Historical playback behind a subscription. | -| [OpenSky Network](https://opensky-network.org/) | Research-oriented, free API, good historical archive. The right choice for bulk work. | -| [Airframes](https://airframes.org/) | Registration-to-airframe history: owners, previous registrations, type. | -| [GPSJam](https://gpsjam.org/) | Daily maps of GPS interference derived from aircraft navigation-accuracy reports. Reveals jamming zones. | +| [OpenSky Network](https://opensky-network.org/) | Research-oriented, documented free API, good historical archive. The right choice for bulk and scripted work. | +| [adsb.lol](https://adsb.lol/) | Community feed with a free, keyless JSON API in the ADS-B Exchange shape. The practical replacement for scripted work now that ADSBx is paid. | +| [Airframes](https://airframes.org/) | Registration-to-airframe history: owners, previous registrations, type, serial. | +| [FAA Registry](https://registry.faa.gov/AircraftInquiry/Search/NNumberInquiry) | Authoritative for US N-numbers: registered owner, address, airworthiness, and the trust structures that obscure them. | +| [GPSJam](https://gpsjam.org/) | Daily maps of GPS interference derived from aircraft navigation-accuracy reports. Reveals jamming zones, which is also why some tracks go wrong. | | [Live ATC](https://www.liveatc.net/) | Archived air-traffic control audio, which sometimes names aircraft that were never tracked. | -OpenSky's API is the one to automate against: +### OpenSky Network API + +The feed to automate against. It publishes state vectors (one position per aircraft per second, +aggregated from its receiver network), per-aircraft flight lists, and full tracks, with a documented +rate model rather than a scraping fight. + +**The auth model changed:** HTTP basic authentication with your account password is gone. Current +access is OAuth2 client credentials — create an API client in your OpenSky account, then exchange +the client ID and secret for a 30-minute bearer token. + +```bash +# one-off: get a token (expires in 30 minutes; a 401 later means refresh it) +TOKEN=$(curl -s -X POST \ + 'https://auth.opensky-network.org/auth/realms/opensky-network/protocol/openid-connect/token' \ + -d 'grant_type=client_credentials' \ + -d 'client_id=YOUR_CLIENT_ID' \ + --data-urlencode 'client_secret=YOUR_CLIENT_SECRET' | jq -r .access_token) +``` + +```bash +# current state of one aircraft by ICAO 24-bit address +curl -s -H "Authorization: Bearer $TOKEN" \ + 'https://opensky-network.org/api/states/all?icao24=4ca7b5' | jq '.states' + +# everything in a bounding box right now — cheaper in credits than a global pull +curl -s -H "Authorization: Bearer $TOKEN" \ + 'https://opensky-network.org/api/states/all?lamin=45&lomin=5&lamax=48&lomax=10' \ + | jq '.states | length' + +# aircraft category as well as position, for sorting military from civil +curl -s -H "Authorization: Bearer $TOKEN" \ + 'https://opensky-network.org/api/states/all?icao24=4ca7b5&extended=1' | jq . + +# state vector at a past instant — the cheap way to answer "where was it at 14:00Z" +# 1773237600 is 2026-03-11T14:00Z; `date -u -d '2026-03-11 14:00' +%s` on GNU, +# `date -u -j -f '%Y-%m-%d %H:%M' '2026-03-11 14:00' +%s` on BSD/macOS +curl -s -H "Authorization: Bearer $TOKEN" \ + 'https://opensky-network.org/api/states/all?icao24=4ca7b5&time=1773237600' | jq . + +# every flight this airframe flew in a two-day window (the maximum interval) +curl -s -H "Authorization: Bearer $TOKEN" \ + 'https://opensky-network.org/api/flights/aircraft?icao24=4ca7b5&begin=1740614400&end=1740787200' \ + | jq '.[] | {firstSeen, estDepartureAirport, lastSeen, estArrivalAirport}' + +# arrivals at an airport, for finding who came in on an unlisted flight +curl -s -H "Authorization: Bearer $TOKEN" \ + 'https://opensky-network.org/api/flights/arrival?airport=EGLL&begin=1740614400&end=1740700800' \ + | jq 'length' + +# the full track of one flight, including the points between airports +curl -s -H "Authorization: Bearer $TOKEN" \ + 'https://opensky-network.org/api/tracks/all?icao24=4ca7b5&time=0' | jq '.path | length' +``` + +Free-tier limits are credit-based and daily: roughly 400 credits anonymous, 4,000 for a registered +account, 8,000 if you feed data back from your own receiver with at least 30% uptime. A +`/states/all` call costs 1–4 credits depending on how large the bounding box is; `/flights/*` and +`/tracks/*` cost far more because they scan day partitions, so a careless historical loop burns a +day's quota in a few calls. `/flights/aircraft` only serves the previous day and earlier — there is +no same-day history. + +What it does not tell you: the hex address identifies a transponder, not an owner, and the +`callsign` field is whatever the crew typed. An aircraft absent from the results was not +necessarily absent from the sky — it may have been out of receiver range, or not transmitting at +all. + +### pyopensky + +The maintained Python client, worth using over hand-rolled requests because it handles the token +refresh and because it also wraps OpenSky's Trino historical database, which the REST API does not +expose. + +```bash +pip install pyopensky +``` + +Credentials go in `settings.conf` (the library prints its location on first run) under a +`[opensky]` section with `client_id` and `client_secret`. + +```python +from pyopensky.rest import REST +from pyopensky.trino import Trino +import datetime as dt + +rest = REST() + +# live state vectors in a bounding box, as a DataFrame +rest.states(bounds=(45, 5, 48, 10)) + +# the route an airline publishes for a callsign — useful when the callsign is the only clue +rest.routes("RYR27LQ") + +# arrivals and departures at an airport over a window +rest.arrival("EGLL", dt.datetime(2026, 3, 11), dt.datetime(2026, 3, 12)) + +# historical flight list for one airframe over months, which REST cannot do +db = Trino() +db.flightlist(start="2026-01-01", stop="2026-03-01", icao24="4ca7b5") + +# the raw trajectory, for plotting a loiter pattern or an orbit +db.history(start="2026-03-11 13:00", stop="2026-03-11 16:00", icao24="4ca7b5") +``` + +Trino access is granted separately from the REST credentials and is intended for academic and +research use — you apply for it, and it is not guaranteed. The REST methods work with ordinary +credentials. Treat Trino as the thing that makes a months-long pattern-of-life analysis possible, +and the REST API as what you use day to day. + +### adsb.lol and ADS-B Exchange + +ADS-B Exchange retired its freemium API tier in March 2025; the current `ADSBexchange.com` RapidAPI +plan starts around $10/month for 10,000 requests, and there is no free key. The map at +[globe.adsbexchange.com](https://globe.adsbexchange.com/) remains free and still shows aircraft +that honour-the-blocklist aggregators hide, so it stays the right place to *look*. For scripted +work, [adsb.lol](https://adsb.lol/) serves the same response shape with no key at all. ```bash -# current state of a specific aircraft by ICAO 24-bit address -curl -s 'https://opensky-network.org/api/states/all?icao24=4ca7b5' | jq . +# one aircraft by ICAO hex — same JSON shape as the ADSBx v2 API +curl -s 'https://api.adsb.lol/v2/hex/4ca7b5' | jq '.ac[] | {flight, r, t, alt_baro, gs, lat, lon}' + +# by callsign, when that is all a photo or an ATC recording gave you +curl -s 'https://api.adsb.lol/v2/callsign/RYR27LQ' | jq '.ac | length' + +# by registration, to confirm a tail number is airborne now +curl -s 'https://api.adsb.lol/v2/reg/EI-EFZ' | jq '.ac[0].hex' + +# every aircraft of a type currently tracked, for fleet questions +curl -s 'https://api.adsb.lol/v2/type/B738' | jq '.total' -# everything in a bounding box right now -curl -s 'https://opensky-network.org/api/states/all?lamin=45&lomin=5&lamax=48&lomax=10' | jq '.states | length' +# everything within 50 nm of a point — the query for "what was over this place" +curl -s 'https://api.adsb.lol/v2/point/51.5/-0.12/50' | jq '.ac[] | {hex, flight, alt_baro}' + +# aircraft squawking an emergency code +curl -s 'https://api.adsb.lol/v2/squawk/7700' | jq '.ac[] | {hex, flight, squawk}' + +# military-tagged aircraft worldwide, right now +curl -s 'https://api.adsb.lol/v2/mil' | jq '.total' + +# LADD and PIA: aircraft using the FAA's blocking and anonymisation programmes +curl -s 'https://api.adsb.lol/v2/ladd' | jq '.total' +curl -s 'https://api.adsb.lol/v2/pia' | jq '.total' ``` -The **ICAO 24-bit hex address** is the durable identifier. Registrations and callsigns change; -the hex usually does not, so track on that. +There is no documented quota, but the service rate-limits aggressively — a burst of requests starts +returning HTTP 429 within seconds, so sleep between calls and cache. Both of these are live-only: +neither gives you history, which is the thing that actually costs money across every ADS-B +provider. + +The `/v2/ladd` and `/v2/pia` endpoints are worth knowing for their own sake. LADD is the FAA +programme that lets an owner suppress their aircraft from public feeds; PIA assigns a temporary +alternative hex address. An aircraft appearing in either list is telling you its owner asked not to +be tracked, which is itself a fact about the aircraft. + +### readsb and dump1090 + +Your own receiver removes the dependency on someone else's coverage, which matters when the area +you care about is a coverage hole. A software-defined radio dongle and an antenna gets you ADS-B +reception out to roughly 200 nautical miles line-of-sight. + +`dump1090` has forked repeatedly; the maintained decoder is +[wiedehopf/readsb](https://github.com/wiedehopf/readsb), with +[tar1090](https://github.com/wiedehopf/tar1090) as the web interface over it. FlightAware's +`dump1090-fa` is still maintained and still fine. Mutability's original `dump1090-mutability` is +abandoned — do not start there. + +```bash +# readsb, the maintained decoder, via the project's install script +sudo bash -c "$(wget -qO - https://github.com/wiedehopf/adsb-scripts/raw/master/readsb-install.sh)" + +# the web interface over it +sudo bash -c "$(wget -qO - https://github.com/wiedehopf/tar1090/raw/master/install.sh)" + +# what your receiver is seeing right now, as JSON +curl -s http://localhost/tar1090/data/aircraft.json | jq '.aircraft | length' + +# log your own feed to disk once a second — this is how you build private history +while true; do + curl -s http://localhost/tar1090/data/aircraft.json >> "feed-$(date -u +%F).jsonl" + echo >> "feed-$(date -u +%F).jsonl" + sleep 1 +done -**Gaps to expect:** ADS-B needs a receiver in range, so oceans, deserts and much of Africa and -central Asia are dark. Military aircraft often transmit nothing. An aircraft that "disappears" -mid-flight has usually just left coverage — check whether the last position is at the edge of a -receiver's range before concluding anything. +# receiver stats, including message rate and position accuracy +curl -s http://localhost/tar1090/data/stats.json | jq '.last1min | {messages, position_rate}' + +# raw Mode S frames on port 30002, for anything the decoder discards +nc localhost 30002 | head -20 + +# Beast binary format on 30005, which is what you feed to aggregators +nc localhost 30005 | xxd | head -5 +``` + +Running your own receiver gives you one thing no aggregator will: the raw message stream with +reception timestamps, which lets you say "my receiver, which was hearing six other aircraft in that +sector at that minute, heard nothing from this one". That is the comparison that turns a gap into +evidence. ## Vessels Ships transmit AIS. Terrestrial receivers reach perhaps 40 nautical miles; beyond that, coverage -depends on satellite AIS, which is mostly commercial. +depends on satellite AIS, which is almost entirely commercial. | Source | Why use it | | --- | --- | | [MarineTraffic](https://www.marinetraffic.com/) | The standard. Live positions, port calls, vessel particulars, photos. Historical track needs a subscription. | -| [VesselFinder](https://www.vesselfinder.com/) | Good free tier; useful cross-check. | +| [VesselFinder](https://www.vesselfinder.com/) | Good free tier; useful cross-check, and its coverage differs from MarineTraffic's in ways that matter when you are arguing about a gap. | | [Equasis](https://www.equasis.org/) | **Ownership and management history, free with registration.** The best free source for who actually controls a ship. | -| [IMO Registry](https://gisis.imo.org/) | Authoritative vessel identity data. | -| [Global Fishing Watch](https://globalfishingwatch.org/map) | Fishing activity inferred from AIS behaviour; exposes probable illegal fishing. | -| [IUU Vessel List](https://iuu-vessels.org/) | Vessels listed for illegal, unreported and unregulated fishing. | +| [IMO GISIS](https://gisis.imo.org/) | Authoritative vessel identity, company and casualty data, free with registration. | +| [Global Fishing Watch](https://globalfishingwatch.org/map) | Fishing effort and vessel-behaviour events inferred from AIS, plus radar-detected vessels that are not broadcasting. | +| [IUU Vessel List](https://iuu-vessels.org/) | Vessels listed by regional fisheries bodies for illegal, unreported and unregulated fishing. | + +### Equasis + +A free, registration-gated database funded by maritime administrations, holding the ownership and +management chain for essentially every merchant ship, plus inspection history. It is the single +most useful free maritime source and it has no API, so this is a walkthrough. + +Register at [equasis.org](https://www.equasis.org/) with a working email; approval is automatic but +can take a day. Then search by IMO number, never by name. + +```text +Ship Info flag, type, tonnage, build year, and the full list of PREVIOUS names + and previous flags with the dates they changed — this is the panel + that exposes a vessel renamed three times in two years +Management Detail registered owner, ISM manager, commercial manager, technical manager, + each with a company IMO number and a country, and each with the date + it took over; the registered owner is usually a one-ship shell, and + the commercial manager is usually who you actually want +Ship History the same chain backwards in time, which is what you cite when an + owner claims they sold the vessel before the incident +Inspections port state control inspections, detentions and the deficiency codes, + which place the ship in a named port on a named date independently + of AIS +``` + +Capture for the record: the company IMO numbers, not just names, because company names repeat +across jurisdictions. Then take each company IMO into +[corporate registries](/sheets/osint/companies-and-finance) — the Equasis company record gives you +a country and an address to start from, and nothing about beneficial ownership. + +Equasis is self-reported by the industry and lags reality by weeks on ownership changes. A +detention record is hard fact; a current owner field is a claim. + +### Global Fishing Watch API + +AIS-derived vessel behaviour as a queryable dataset rather than a map: identity resolution across +MMSI and IMO, and coded events — encounters between two vessels, loitering, port visits, and AIS +gaps. The gap detection is the part that matters outside fisheries work, because it has already +done the coverage reasoning for you. -The **IMO number** is the durable identifier — it stays with the hull for life. Names, flags and -owners change constantly, often specifically to frustrate tracking, so always record the IMO. +Get a token at [globalfishingwatch.org/our-apis/tokens](https://globalfishingwatch.org/our-apis/tokens). +Access is free and non-commercial only. -**AIS is deliberately gamed.** Transponders get switched off during transfers, and positions are -sometimes spoofed outright. A vessel going dark for six hours near another dark vessel is a -finding, not a data problem. +```bash +TOKEN='YOUR_GFW_TOKEN' + +# resolve an identifier to GFW's vessel id, with flag and name history +curl -s -H "Authorization: Bearer $TOKEN" -G \ + 'https://gateway.api.globalfishingwatch.org/v3/vessels/search' \ + --data-urlencode 'query=9876543' \ + --data-urlencode 'datasets[0]=public-global-vessel-identity:latest' \ + | jq '.entries[] | {id:.selfReportedInfo[0].id, name:.selfReportedInfo[0].shipname, flag:.selfReportedInfo[0].flag}' + +# encounters: two vessels close and slow for long enough to transfer something +curl -s -H "Authorization: Bearer $TOKEN" -G \ + 'https://gateway.api.globalfishingwatch.org/v3/events' \ + --data-urlencode 'datasets[0]=public-global-encounters-events:latest' \ + --data-urlencode 'vessels[0]=VESSEL_ID' \ + --data-urlencode 'start-date=2026-01-01' --data-urlencode 'end-date=2026-03-01' | jq '.total' + +# AIS gaps — transmission stopping and resuming, with both endpoints +curl -s -H "Authorization: Bearer $TOKEN" -G \ + 'https://gateway.api.globalfishingwatch.org/v3/events' \ + --data-urlencode 'datasets[0]=public-global-gaps-events:latest' \ + --data-urlencode 'vessels[0]=VESSEL_ID' \ + --data-urlencode 'start-date=2026-01-01' --data-urlencode 'end-date=2026-03-01' \ + | jq '.entries[] | {start, end, lat:.position.lat, lon:.position.lon}' + +# loitering: slow movement away from port, often the other half of an encounter +curl -s -H "Authorization: Bearer $TOKEN" -G \ + 'https://gateway.api.globalfishingwatch.org/v3/events' \ + --data-urlencode 'datasets[0]=public-global-loitering-events:latest' \ + --data-urlencode 'vessels[0]=VESSEL_ID' \ + --data-urlencode 'start-date=2026-01-01' --data-urlencode 'end-date=2026-03-01' | jq '.total' + +# port visits, which date the vessel to a jurisdiction +curl -s -H "Authorization: Bearer $TOKEN" -G \ + 'https://gateway.api.globalfishingwatch.org/v3/events' \ + --data-urlencode 'datasets[0]=public-global-port-visits-events:latest' \ + --data-urlencode 'vessels[0]=VESSEL_ID' \ + --data-urlencode 'start-date=2026-01-01' --data-urlencode 'end-date=2026-03-01' | jq '.total' +``` + +A GFW gap event is a reasoned inference, not a raw observation: it accounts for expected satellite +coverage and known reception quality in that cell, which is exactly the work you would otherwise +have to do yourself. It is still an inference — a receiver outage or an equipment failure produces +the same signature as a deliberate switch-off. + +## Evidencing a gap + +The difference between "went dark" as a finding and as a guess is whether you can show the data +should have been there. + +```text +1. Fix the window last position with a timestamp, first position after, both with + the reporting source named +2. Establish coverage other aircraft or vessels reported in the same cell during the + same window, from the same feed. If nothing else reported either, + you have a coverage hole, not a dark period +3. Cross-feed check a second, independent network: adsb.lol against OpenSky, + VesselFinder against MarineTraffic. Independent receiver + populations failing identically is coverage; one seeing traffic + while the other does not is a reception artefact +4. Check the physics the distance between last and first position divided by the gap + duration gives an implied speed. An implied 60 knots for a bulk + carrier means the track is wrong, not that the ship was fast +5. Look for the reason GPSJam for the aircraft case, since jamming degrades position + quality before it removes it; a port call or an AIS equipment + deficiency in the Equasis inspection record for the vessel case +6. Corroborate satellite imagery over the implied position during the gap. + A radar or optical detection of a hull where AIS says nothing + is the strongest version of this finding +``` + +The three mechanisms to distinguish between: **transponder off**, where transmission simply stops +and resumes, usually near a transfer or a sanctioned port; **MMSI spoofing**, where the identifier +is set to another vessel's or an invalid one, detectable as two vessels with one MMSI in different +oceans, or an MMSI whose country prefix contradicts the claimed flag; and **position spoofing**, +where the track is plausible but false, detectable as an implied speed the hull cannot do, a track +that crosses land, or imagery showing the berth empty. ## Rail and road - [OpenRailwayMap](https://www.openrailwaymap.org/) — global rail infrastructure, including - electrification, gauge and signalling. -- [Chronotrains](https://www.chronotrains.com/) — how far you can travel by train in N hours. + electrification, gauge and signalling, which constrains what can run on a line. +- [Chronotrains](https://www.chronotrains.com/) — how far you can travel by train in N hours, + useful for bounding where someone could have been. - [License Plate Maps](https://www.licenseplatemania.com/) — plate formats by country, for - narrowing a location from a vehicle. + narrowing a location from a vehicle in a photograph. +- [VIN Decoder (NHTSA)](https://vpic.nhtsa.dot.gov/decoder/) — free US government VIN decode, with + a JSON API: `curl -s 'https://vpic.nhtsa.dot.gov/api/vehicles/decodevin/1HGCM82633A004352?format=json'` ## Tool reference @@ -112,6 +433,75 @@ finding, not a data problem. you need when you see it. - **Flags of convenience** mean the flag state tells you little about real ownership. Use Equasis. +## Worked example + +You have one datum: the tail number `N721AB`, from a photograph of an aircraft on a stand. The +registration, hex and outputs below are illustrative — the sequence of pivots is the point. + +```bash +# 1. registration to hex, from a live feed — the hex is what you track on +curl -s 'https://api.adsb.lol/v2/reg/N721AB' | jq '.ac[0] | {hex, t, flight, alt_baro}' +# {"hex":"a8a2b5","t":"GLF6","flight":null,"alt_baro":"ground"} +``` + +The airframe is a Gulfstream G650ER and it is on the ground somewhere right now. The hex `a8a2b5` +is the identifier for everything that follows. + +```bash +# 2. registered owner, from the FAA registry — not the operator +# registry.faa.gov/AircraftInquiry/Search/NNumberInquiry, N721AB +# -> registered to a Delaware LLC, c/o a Nevada registered agent +``` + +A single-purpose LLC behind a registered agent is the normal arrangement for a private jet, so the +registry gives you a company name, not a person. That name goes to +[corporate records](/sheets/osint/companies-and-finance). + +```bash +# 3. the movement history, from OpenSky (begin/end are 2025-02-27 to 03-01 — +# two days is the maximum interval /flights/aircraft accepts) +TOKEN=$(curl -s -X POST \ + 'https://auth.opensky-network.org/auth/realms/opensky-network/protocol/openid-connect/token' \ + -d 'grant_type=client_credentials' -d 'client_id=ID' \ + --data-urlencode 'client_secret=SECRET' | jq -r .access_token) + +curl -s -H "Authorization: Bearer $TOKEN" \ + 'https://opensky-network.org/api/flights/aircraft?icao24=a8a2b5&begin=1740614400&end=1740787200' \ + | jq '.[] | {dep:.estDepartureAirport, arr:.estArrivalAirport, off:.firstSeen, on:.lastSeen}' +# three legs across the window; the second shows estDepartureAirport "KVNY" +# and estArrivalAirport null +``` + +A leg with a departure and no arrival is the pivot. Either the aircraft left coverage, or the +transponder stopped. + +```bash +# 4. the raw track for that leg, to find where the positions stop +curl -s -H "Authorization: Bearer $TOKEN" \ + 'https://opensky-network.org/api/tracks/all?icao24=a8a2b5&time=1740671000' \ + | jq '.path[-3:]' +# last point 31.2N 118.4W at FL410, over open water south-west of San Diego +``` + +```bash +# 5. was that a coverage hole? ask what else reported in the same cell +curl -s -H "Authorization: Bearer $TOKEN" \ + 'https://opensky-network.org/api/states/all?time=1740671400&lamin=30&lomin=-120&lamax=33&lomax=-116' \ + | jq '[.states[] | {icao:.[0], alt:.[13]}] | length' +# 0 — nothing at all reported in that box at that minute +``` + +Zero other aircraft in a box that size over a busy approach corridor is a coverage answer, not an +evasion answer: the aircraft flew out of the receiver network's reach over the Pacific. Checking +[adsb.lol](https://adsb.lol/) for the same window returns nothing either, from an independent +receiver population — which confirms coverage rather than contradicting it. + +What the chain gives you: a type, a hex, a registered owner to take into company records, three +dated legs, and a documented reason to *not* report a dark period. What it does not give you is who +was on board, which no transponder feed ever will — that needs a stand photograph with a timestamp, +a flight-plan filing, or a planespotter's log, and the second independent source is what would make +it defensible. + ## Broader catalogues - [Transport OSINT](https://tools.osintnewsletter.com/tool-categories/transport-osint) diff --git a/src/content/sheets/osint/websites-and-infrastructure.md b/src/content/sheets/osint/websites-and-infrastructure.md @@ -4,7 +4,7 @@ description: "Attribute a site: registration history, DNS, certificates, analyti category: osint subcategory: "Infrastructure" tags: [osint, domains, dns, certificates, attribution] -tools: [urlscan, crtsh, shodan, wayback] +tools: [urlscan, crtsh, shodan, subfinder, httpx, wayback, rdap] difficulty: intermediate updated: 2026-09-28 references: @@ -44,84 +44,378 @@ surface. 6. **Check what shares the host.** Shared hosting is meaningless; a dedicated IP hosting five related domains is not. -## Registration and DNS +## Page fingerprinting + +Linking two sites to one operator comes down to finding a value that is identical in both and that +nobody would share by accident. The signals are not equal: + +| Signal | Why it links sites | Strength | +| --- | --- | --- | +| Google Analytics / Tag Manager ID | Operators reuse one property across their sites constantly. | Near-conclusive | +| AdSense publisher ID | Same, and tied to a payment account. | Near-conclusive | +| TLS certificate serial or key reuse | The same certificate on two hosts means one operator. | Strong | +| Favicon hash | Shodan indexes it, so one favicon finds every host serving it. | Strong if distinctive | +| Distinctive HTML comment or typo | Copy-pasted templates carry unique strings. | Strong | +| Response-body hash | Identical pages on different hosts. | Moderate | +| Shared IP on dedicated hosting | Meaningful only on a small host, never on a CDN. | Weak to worthless | +| Same registrar or nameserver | Millions of unrelated domains share these. | Worthless alone | + +Searching an ID string in [PublicWWW](https://publicwww.com/) or [Grep.app](https://grep.app/) +finds the other pages that embed it; `urlscan.io` finds the sites whose scans *requested* it, which +catches IDs injected at runtime rather than written into the HTML. + +## Key tools + +### RDAP and whois + +RDAP is the structured replacement for `whois`, and the one to reach for first: JSON instead of +free text, consistent field names across registries, and no scraping. `whois` is still worth +running because some registries put things in the free-text blob that never make it into RDAP. ```bash -# current registration -whois example.com +# RDAP, following the bootstrap redirect to whichever registry is authoritative +curl -s -L 'https://rdap.org/domain/example.com' | jq '{ldhName,status,events,entities}' -# DNS basics -dig +short A example.com -dig +short MX example.com -dig +short TXT example.com -dig +short NS example.com +# the registration and expiry dates on their own +curl -s -L 'https://rdap.org/domain/example.com' \ + | jq -r '.events[] | "\(.eventAction) \(.eventDate)"' + +# nameserver history is not here, but the current delegation is +curl -s -L 'https://rdap.org/domain/example.com' | jq -r '.nameservers[].ldhName' + +# straight to the registry when you know it, which is faster and never rate-limited by a proxy +curl -s 'https://rdap.verisign.com/com/v1/domain/example.com' | jq '.events' + +# an IP or a netblock: who holds the allocation, and when it was assigned +curl -s -L 'https://rdap.org/ip/93.184.216.34' | jq '{handle,name,country,events}' + +# an ASN, for working out whose network a host actually sits on +curl -s -L 'https://rdap.org/autnum/15169' | jq '{handle,name,entities}' + +# the old tool, for the fields RDAP drops +whois example.com | grep -iE 'registrar|created|updated|expir|name server|status' +``` + +Current records are privacy-shielded almost everywhere, so RDAP mostly gives you dates, status +codes and the registrar — all useful, none of them a name. The creation date is the one field +worth trusting: it is set once and does not change, which makes it the backbone of any timeline. +Registrar-level and ccTLD RDAP coverage is patchy; a 404 from `rdap.org` means no server is +registered for that zone, not that the domain is free. + +### dig and passive DNS + +Live DNS tells you the present; passive DNS — a third party's historical record of what resolved +to what — tells you the past, which is where attribution lives. Run both: the live records name +the vendors in use, and the history names the host before the CDN went up. + +```bash +# the basics, worth doing as a set rather than one at a time +for r in A AAAA MX NS TXT SOA CNAME; do echo "== $r"; dig +short "$r" example.com; done -# mail and verification records often name the vendors in use +# mail and verification records enumerate the SaaS vendors the operator uses +dig +short TXT example.com dig +short TXT _dmarc.example.com +dig +short TXT _domainkey.example.com + +# ask the authoritative server directly, bypassing your resolver's cache +dig @"$(dig +short NS example.com | head -1)" example.com ANY + +# reverse lookup on the IP, which occasionally names the hosting account +dig +short -x 93.184.216.34 + +# the full delegation chain, for spotting a subdomain handed to a third party +dig +trace staging.example.com | tail -20 + +# a zone transfer, which should fail — when it does not, you have the whole zone +dig @ns1.example.com example.com AXFR ``` -For history, use [DNS History](https://dnshistory.org/), Whoxy or a Domain Research Suite — the -value is in what changed and when, not the current state. +`dig +short NS` plus `AXFR` is the one active step here; everything else above only queries +resolvers. For history, [DNS History](https://dnshistory.org/) and [Whoxy](https://www.whoxy.com/) +are the usual free-tier sources, and [SecurityTrails](https://securitytrails.com/) has the deepest +free passive-DNS allowance, metered per month. A shared IP in the history is worthless; a +dedicated IP that three domains shared for two years is a finding. -## Certificate transparency +### crt.sh -Every publicly trusted certificate is logged, so CT logs are a free and complete subdomain -enumerator: +Certificate transparency logs are public, complete and free, which makes them the best subdomain +enumerator in existence — and it never touches the target. crt.sh is the searchable front end, +and its JSON output is the part to automate. ```bash -# all names ever certificated for a domain +# every name ever certificated under the domain, deduplicated curl -s 'https://crt.sh/?q=%25.example.com&output=json' \ - | jq -r '.[].name_value' | tr '\n' '\n' | sort -u + | jq -r '.[].name_value' | tr ' ' '\n' | sed 's/^\*\.//' | sort -u -# just the certificate issuance timeline +# with the issuance timeline, which dates when each host appeared curl -s 'https://crt.sh/?q=example.com&output=json' \ | jq -r '.[] | "\(.not_before) \(.name_value)"' | sort -u | head -40 + +# who issues their certificates — a change of CA often marks a change of operator +curl -s 'https://crt.sh/?q=%25.example.com&output=json' \ + | jq -r '.[].issuer_name' | sort | uniq -c | sort -rn + +# search by organisation name instead of domain, to find their other estates +curl -s 'https://crt.sh/?q=Example+Organisation+Ltd&output=json' | jq -r '.[].common_name' | sort -u + +# one certificate's full detail by crt.sh id +curl -s 'https://crt.sh/?id=29510272303&output=json' | jq '{issuer_name,not_before,not_after,serial_number}' + +# only names that still resolve, so you can tell live infrastructure from history +curl -s 'https://crt.sh/?q=%25.example.com&output=json' | jq -r '.[].name_value' \ + | tr ' ' '\n' | sed 's/^\*\.//' | sort -u \ + | while read -r h; do [ -n "$(dig +short "$h")" ] && echo "$h"; done ``` -A certificate issued for a hostname that does not resolve publicly still tells you the hostname -exists. +crt.sh is run by Sectigo as a public service and is frequently overloaded — a 502 on the web UI is +normal and usually clears, and the JSON endpoint stays up more often than the HTML one does. When +it is down, `subfinder` pulls the same logs through other indexes. A certificate proves a name was +requested, not that a host ever answered on it: staging and internal hostnames show up here +precisely because nobody expected them to be published. -## Page fingerprinting +### subfinder -| Signal | Why it links sites | -| --- | --- | -| Google Analytics / Tag Manager ID | Operators reuse one property across their sites constantly. The single strongest link. | -| AdSense publisher ID | Same. | -| Favicon hash | Shodan indexes it, so one favicon finds every host serving it. | -| Distinctive HTML comment or typo | Copy-pasted templates carry unique strings. | -| TLS certificate serial / key reuse | Same cert on two hosts means one operator. | +Passive subdomain enumeration across several dozen indexes at once — certificate logs, passive +DNS providers, search engines, threat-intel feeds. Faster and broader than querying crt.sh alone, +and it needs no keys to be useful, though adding free ones roughly doubles what it finds. + +```bash +# install: a single Go binary +go install -v github.com/projectdiscovery/subfinder/v2/cmd/subfinder@latest + +# the default run: passive only, every keyless source +subfinder -d example.com -silent + +# which sources exist, and which ones your keys unlock +subfinder -ls + +# all sources including the slow ones, written to a file +subfinder -d example.com -all -o subs.txt + +# JSON lines with the source recorded per host, so a finding is attributable +subfinder -d example.com -oJ -cs -o subs.jsonl + +# several root domains in one run +subfinder -dL roots.txt -oD ./out -silent -[The Information Laundromat](https://information-laundromat.com/) automates much of this, -comparing content and technical indicators across sites to surface shared operators. -[Urlscan](https://urlscan.io/) records what a page loaded, including third-party requests, and its -archive is searchable — so you can find other sites that made the same unusual request. +# only hosts that actually resolve, with their IPs +subfinder -d example.com -nW -oI -silent + +# throttle a chatty provider rather than getting your free key banned +subfinder -d example.com -rls 'crtsh=1/s' -rl 10 -silent +``` + +Free API keys go in the provider config under your platform's config directory; `subfinder -ls` +shows which sources are active. Passive results are historical, so a long list includes hosts +retired years ago — pass `-nW` when you need the live set. Sources disagree and some return +wildcard noise, so a single-source hit deserves `-cs` output and a second look before you build on +it. + +### urlscan.io + +A hosted browser that visits a page for you, records every request it made, and keeps the result +in a searchable archive. That gives you two separate capabilities: look at a hostile page without +touching it yourself, and search what *other* people's scans loaded to find sites sharing an +unusual third-party request. ```bash -# search urlscan's archive without visiting the target -curl -s 'https://urlscan.io/api/v1/search/?q=page.domain%3Aexample.com' \ - | jq -r '.results[] | "\(.task.time) \(.page.url)"' | head +# search the archive without visiting anything — no key needed for modest volume +curl -s 'https://urlscan.io/api/v1/search/?q=domain%3Aexample.com&size=20' \ + | jq -r '.results[] | "\(.task.time) \(.page.url)"' + +# every site whose scans loaded a specific analytics ID: the strongest operator link there is +curl -s 'https://urlscan.io/api/v1/search/?q=page.url%3A*+AND+%22UA-12345678%22' | jq '.total' + +# scans by the IP a page resolved to, for finding co-hosted estates +curl -s 'https://urlscan.io/api/v1/search/?q=page.ip%3A93.184.216.34&size=50' \ + | jq -r '.results[].page.domain' | sort -u + +# ASN-wide search, which is how you sweep a bulletproof host +curl -s 'https://urlscan.io/api/v1/search/?q=page.asn%3AAS15169+AND+page.domain%3A*.example' | jq '.total' + +# submit a scan — the header name is API-Key, and nothing else works +curl -s -X POST 'https://urlscan.io/api/v1/scan/' \ + -H "API-Key: $URLSCAN_KEY" -H 'Content-Type: application/json' \ + -d '{"url":"https://example.com","visibility":"unlisted"}' | jq -r '.uuid' + +# then read the result a minute later: requests, certificates, redirect chain, screenshot +curl -s 'https://urlscan.io/api/v1/result/UUID/' \ + | jq '{page:.page, redirects:[.data.requests[].request.redirectResponse.url]}' -# find pages sharing a specific analytics ID -curl -s 'https://urlscan.io/api/v1/search/?q=page.url%3A*%20AND%20UA-12345678' | jq '.total' +# your remaining quota, which is the only reliable statement of your limits +curl -s 'https://urlscan.io/user/quotas/' -H "API-Key: $URLSCAN_KEY" ``` -Searching a code-search engine such as [Grep.app](https://grep.app/) or -[PublicWWW](https://publicwww.com/) for an analytics ID finds the other sites that embed it. +`visibility` is the decision that matters. A `public` scan is visible to everyone immediately, +including your target, who may well be watching the archive for their own domains; `unlisted` +keeps it out of search but still reachable by URL. Unauthenticated search quotas are small and +reset on the full minute and hour, so batch your queries. A scan is one browser, from a datacentre +IP, in one country — a page that cloaks on geography or user agent will show you something the +target's visitors never see. -## Archives +### Shodan -The [Wayback Machine](https://web.archive.org/) is the primary source. Its CDX API is the part -worth automating: +The index of what is listening on the internet, and the only tool here that can answer "what else +serves this exact favicon". Its value in attribution work is the pivot: a hash, a certificate +serial or an HTTP header becomes a list of hosts. ```bash -# every capture of a path, with timestamps and status -curl -s 'http://web.archive.org/cdx/search/cdx?url=example.com/contact*&output=json&collapse=digest' +# install and register the key once +pipx install shodan +shodan init YOUR_API_KEY -# first and last capture of a domain -curl -s 'http://web.archive.org/cdx/search/cdx?url=example.com&output=json&limit=1' +# free-tier reality check: how many results a query has, before spending credits on it +shodan count 'http.favicon.hash:-247388890' + +# the pivot that matters — every host serving the same favicon +shodan search --fields ip_str,port,hostnames,org 'http.favicon.hash:-247388890' + +# everything known about one address, including historical ports +shodan host 93.184.216.34 + +# hosts sharing a TLS certificate serial, i.e. the same operator +shodan search --fields ip_str,hostnames 'ssl.cert.serial:"73d94ae1858e6f4613fd389a8efe5267"' + +# a distinctive response header or body string, which copy-pasted templates carry +shodan search --fields ip_str,http.title 'http.html:"Example Org internal portal"' + +# facet counts, for describing an estate rather than listing it +shodan stats --facets country:10,org:10 'ssl.cert.subject.CN:*.example.com' + +# the subdomains and DNS records Shodan has seen for a domain +shodan domain example.com + +# bulk download then parse offline, so you pay for the query once +shodan download results.json.gz 'http.favicon.hash:-247388890' +shodan parse --fields ip_str,port,org --separator , results.json.gz ``` -`archive.today` catches things Wayback misses and is harder to have removed. +Search filters need a paid or academic key; a free account can look up hosts but not run filtered +queries, so budget for that before planning a workflow around it. Installed under Python 3.12 the +CLI can fail with `ModuleNotFoundError: pkg_resources` — install `setuptools` into the same +environment. Shodan's data is a crawl with a lag of days to weeks, so a result describes what was +listening at scan time, and `hostnames` are reverse-DNS, not proof the site is served there. + +### httpx + +Probes a list of hosts and reports what answered, with the fingerprints you need for linking sites +in one pass: favicon hash, JARM, title, technology, body hash. The natural next step after +`subfinder`, and the way you generate the favicon hash that Shodan then searches. + +```bash +# install: ProjectDiscovery's httpx, NOT the Python HTTP library of the same name +go install -v github.com/projectdiscovery/httpx/cmd/httpx@latest + +# what is actually live, with status and title +subfinder -d example.com -silent | httpx -silent -sc -title + +# the fingerprint set, as JSON you can diff between estates +httpx -l subs.txt -json -favicon -jarm -tech-detect -title -web-server -o fingerprints.jsonl + +# the favicon hash on its own — this is the number you paste into Shodan +httpx -u https://example.com -favicon -silent + +# group hosts by response-body hash to find the ones serving identical pages +httpx -l subs.txt -hash mmh3 -silent | sort -k2 | uniq -c -f1 | sort -rn | head + +# whether a CDN is hiding the origin, and whose network it is +httpx -l subs.txt -silent -cdn -ip -asn + +# screenshots, for the record rather than the analysis +httpx -l subs.txt -screenshot -silent -o shots.txt + +# slow it down: this is the one tool on this page that touches the target +httpx -l subs.txt -rate-limit 5 -silent -sc +``` + +This is active: every probe is a request from your address to their server, logged at their end. +Everything above it on this page is passive, so do the passive work first and know what you are +looking for before you knock. A favicon hash is only a link when the favicon is distinctive — +thousands of hosts share the default WordPress and cPanel icons, so check the hash's global count +in Shodan before treating a match as meaningful. + +### Wayback Machine CDX API + +The archive's index, and far more useful than the calendar view: it lists every capture of a URL +pattern with timestamp, status code and content digest, which turns "what did this site used to +say" into a diffable dataset. + +```bash +CDX='https://web.archive.org/cdx/search/cdx' + +# every distinct capture of a path, collapsed so identical content appears once +curl -s -G "$CDX" --data-urlencode 'url=example.com/contact*' \ + --data-urlencode 'output=json' --data-urlencode 'collapse=digest' + +# the first and last capture, i.e. when the site appeared and when it stopped +curl -s -G "$CDX" --data-urlencode 'url=example.com' --data-urlencode 'output=json' \ + --data-urlencode 'limit=1' +curl -s -G "$CDX" --data-urlencode 'url=example.com' --data-urlencode 'output=json' \ + --data-urlencode 'limit=-1' + +# every URL ever archived under the domain — the archive as a site map +curl -s -G "$CDX" --data-urlencode 'url=example.com/*' --data-urlencode 'output=json' \ + --data-urlencode 'fl=original' --data-urlencode 'collapse=urlkey' | jq -r '.[1:][] | .[0]' + +# only captures that returned 200, within a date window +curl -s -G "$CDX" --data-urlencode 'url=example.com/*' --data-urlencode 'output=json' \ + --data-urlencode 'filter=statuscode:200' --data-urlencode 'from=2018' --data-urlencode 'to=2020' + +# pull one archived page as it was served, without the archive's own toolbar +curl -s 'https://web.archive.org/web/20190101120000id_/http://example.com/contact' + +# diff two captures of the staff page, which is where removed names live +for ts in 20180301000000 20210301000000; do + curl -s "https://web.archive.org/web/${ts}id_/http://example.com/team" > "team-$ts.html" +done +diff team-20180301000000.html team-20210301000000.html | head -40 +``` + +Use `https` — the `http` form of the CDX endpoint returns nothing rather than redirecting. The +`id_` suffix on a snapshot URL serves the original bytes, which is what you want for forensics and +for hashing. The archive honours removal requests and `robots.txt` retroactively in places, so +absence is not evidence of absence; [archive.today](https://archive.ph/) catches things Wayback +missed and is harder to have taken down. See +[archiving and evidence](/sheets/osint/archiving-and-evidence) for making your own captures stand +up. + +### The Information Laundromat + +Web only, free, and purpose-built for the question the rest of this page answers one signal at a +time: are these sites run by the same people. Paste in a URL or a block of content and it compares +both the text and the technical indicators across its index. + +Use it at [informationlaundromat.com](https://informationlaundromat.com/) — note the hyphenated +spelling of the domain no longer resolves. Two modes matter: + +```text +Content similarity paste an article; it finds the other sites republishing the same text, + which is how a syndication network becomes visible +Metadata similarity give it a domain; it compares analytics IDs, ad IDs, CDN and registrar + fingerprints against its index and scores the overlap +Read the result as: which indicator matched, not the headline score — a shared Google + Analytics ID is near-conclusive, a shared Cloudflare nameserver is noise +Capture: the matched indicator, its value, and the date you ran it +``` + +The scoring blends strong and weak indicators, so a high score built entirely on shared hosting +means nothing. Open the per-indicator breakdown every time, and confirm the strong ones by hand — +a claimed analytics ID should be visible in the page source or in a +[PublicWWW](https://publicwww.com/) or [Grep.app](https://grep.app/) search before you rely on it. + +### Tools that have moved since this area settled + +- **Censys** is retiring Legacy Search (`search.censys.io`) and its v1/v2 APIs through 2026, in + favour of the Platform API at `https://api.platform.censys.io/v3/global/`. Scripts written + against `/api/v2/hosts/...` need rewriting; certificate observation endpoints have already + dropped from realtime to a six-hourly refresh. +- **Amass** is at v5 and no longer works the way most cheatsheets describe. Passive is now the + default, so `-passive` is a deprecated no-op and `-active` is the flag that adds zone transfers + and certificate grabs. Results land in its own asset database, read back with + `amass subs` and `amass viz` rather than printed by `amass enum`. For plain passive subdomain + lists, `subfinder` is the simpler tool. +- **DNSDumpster** now requires an account for anything beyond a handful of lookups a day. The + keyless equivalents are crt.sh and `subfinder`. ## Tool reference @@ -156,6 +450,42 @@ curl -s 'http://web.archive.org/cdx/search/cdx?url=example.com&output=json&limit not. - **Archives have gaps and honour exclusions.** Absence from Wayback is not absence from the web. +## Worked example + +One datum: the domain **example-news-daily.com**, which published a story you are checking. The +question is who runs it and what else they run. + +1. **Dates first.** `curl -s -L 'https://rdap.org/domain/example-news-daily.com'` gives a + registration event in March 2024 and a privacy-shielded registrant. A site presenting itself as + a long-running local paper, registered eighteen months ago, is already a finding. +2. **Names it has had.** crt.sh for `%.example-news-daily.com` returns six hostnames, including + `old.` and `wp.` subdomains and — more usefully — a certificate from 2024 that also covers + `example-city-times.com`. One certificate across two brands is a strong link. +3. **Broaden the host list.** `subfinder -d example-news-daily.com -all -oJ -cs` adds four names + crt.sh missed and records which source found each, so you can cite them individually. +4. **See what is live without being seen.** A urlscan archive search for + `domain:example-news-daily.com` returns existing public scans, including the third-party requests the page makes. One is a Google + Analytics beacon carrying a property ID. +5. **Pivot on the ID.** That ID searched in PublicWWW returns eleven other domains embedding it, + and the urlscan query `page.url:* AND "UA-…"` returns nine of the same. Overlapping lists from + two independent indexes is what makes the network claim defensible. +6. **Confirm the shared template.** `httpx -l hosts.txt -json -favicon -jarm -title` shows the same + favicon hash across the whole set. That hash in `shodan count` returns 14 hosts globally — small + enough that the match means something, which a default-WordPress icon would not. +7. **Read the history.** The Wayback CDX index for `example-news-daily.com/*` shows the earliest + capture in April 2024 and an `/about` page archived in May 2024 naming two editors; the live + `/about` names neither. `diff` of the two captures is the record. +8. **Check the network, not just the pair.** The Information Laundromat's metadata comparison on the + domain scores the same cluster and surfaces two more domains that share the ad ID but not the + analytics ID — a wider ring, confirmed on a second indicator. + +What you can assert: a domain registered in March 2024, sharing a certificate, an analytics +property, an ad publisher ID and a distinctive favicon with eleven other sites, with two editor +names removed from its own about page in 2024 and preserved in the archive. What you cannot: who +those people are. Nothing on this page crosses from infrastructure to identity — that pivot runs +through [people search](/sheets/osint/people-search) and +[email and phone work](/sheets/osint/email-and-phone). + ## Broader catalogues - [Domain Name OSINT](https://tools.osintnewsletter.com/tool-categories/domain-name-osint)