daemon-sec-cheatsheet

The cheatsheet vault for operators: AD, enumeration, exploitation, priv-esc, web, DFIR
git clone https://git.daemon-sec.xyz/daemon-sec-cheatsheet.git
Log | Files | Refs | README | LICENSE

commit d1f3335bc7a7af813e54950b0f8b0027596f0317
parent 578fa3b5dc8ce50fd6216a3cc255e55e6b2ac9c0
Author: DAEMON <zer0sec.xp@icloud.com>
Date:   Sat,  3 Oct 2026 18:51:29 +0100

Finish the OSINT usage pass: the last six sheets

Completes the OSINT tool-usage work. archiving-and-evidence,
data-analysis-and-visualisation, osint-foundations, geolocation,
reverse-image-search and social-media-monitoring were the remaining
catalogues — between 0 and 3 code blocks each, no per-tool sections.
They now follow usernames-and-accounts: a numbered Method, per-tool
sections with install and six to ten commented invocations, a closing
note on limits, and a worked example from one concrete datum. 14 of 16
OSINT sheets done; the two untouched were already the reference shape.

Flags and endpoints were run or probed rather than recalled, which again
found breakage in what was already published:

- information-laundromat.com is NXDOMAIN; the live tool is
  informationlaundromat.com with no hyphen.
- Anonymous Wayback GET /save/<url> returns 429 with a body that looks
  like a page, so old scripts fail silently. Replaced with SPN2.
- Bing's Visual Search API was decommissioned in August 2025.
- Public rsshub.app returns 403, and bridge routes fail silently with an
  empty feed rather than an error.
- ShadowFinder's CLI is positional under a subcommand now.
- csvkit type inference rewrites 01/03/2026 to 2026-01-03 silently, and
  pandas format='mixed' guards against silent NaT coercion.
- wpull is dead and grab-site stale, so browsertrix-crawler is
  documented for WARC/WACZ instead.

Also corrects image-video-forensics, which claimed -vsync vfr and
-fps_mode vfr are both accepted: -vsync was removed outright in ffmpeg 9
and fails with 'Unrecognized option'. And trims the /credits ired.team
blurb, which claimed derived sheets in three categories that have none.

The geolocation worked example's two-date-windows claim was checked by
sweeping a full year at five-minute resolution rather than asserted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Diffstat:
Msrc/content/sheets/osint/archiving-and-evidence.md | 682++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-----
Msrc/content/sheets/osint/data-analysis-and-visualisation.md | 625++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-------
Msrc/content/sheets/osint/geolocation.md | 457++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-------
Msrc/content/sheets/osint/image-video-forensics.md | 5+++--
Msrc/content/sheets/osint/osint-foundations.md | 596++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++---
Msrc/content/sheets/osint/reverse-image-search.md | 377+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++----------
Msrc/content/sheets/osint/social-media-monitoring.md | 508+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++----
Msrc/pages/credits.astro | 5++---
8 files changed, 3030 insertions(+), 225 deletions(-)

diff --git a/src/content/sheets/osint/archiving-and-evidence.md b/src/content/sheets/osint/archiving-and-evidence.md @@ -4,9 +4,9 @@ description: "Capture sources so they survive deletion, with the hashes and time category: osint subcategory: "Archiving & Analysis" tags: [osint, archiving, evidence, preservation] -tools: [auto-archiver, wayback, archive-today, yt-dlp] -difficulty: beginner -updated: 2026-09-28 +tools: [wget, browsertrix-crawler, yt-dlp, auto-archiver, wayback, archive-today, opentimestamps, openssl, exiftool, shasum] +difficulty: intermediate +updated: 2026-10-04 references: - name: "Bellingcat's Online Investigation Toolkit" url: "https://bellingcat.gitbook.io/toolkit" @@ -24,56 +24,547 @@ references: ## What this covers -Making sure what you found still exists tomorrow, and that you can show it has not changed since -you found it. Content gets deleted, edited and made private constantly — usually right after -someone notices attention. Archive first, analyse second. +Capturing a source so that it still exists after someone deletes it, and so that you can show the +copy you hold is the copy you took. The capture itself is the easy half; the hash, the third-party +timestamp and the written provenance are what make it worth anything to an editor, a lawyer or a +court. Everything below is organised around one question: what does this step actually prove? ## Method -1. **Archive to something you do not control** — Wayback, archive.today — so the capture has a - third-party timestamp. -2. **Also keep a local copy.** Public archives can be removed, rate-limited or unavailable. -3. **Hash every local file** on acquisition, and record the hash somewhere separate. -4. **Record provenance:** the URL, the moment you captured it, and how you reached it. -5. **Capture the media, not a screenshot of it.** A screenshot loses metadata, resolution and - audio. +1. **Check what already exists before you capture.** A Wayback snapshot from before the content + became interesting is worth more than one you took after. Query the availability and CDX APIs + first; if there is an older capture, that is your baseline and the diff against today is itself a + finding. +2. **Capture to something you do not control, and to something you do.** A third-party archive + supplies a timestamp nobody can accuse you of forging. A local copy survives the archive being + rate-limited, robots-excluded or taken down. Neither substitutes for the other. +3. **Pick the capture tool by what the page is.** Static HTML and media: `wget --warc-file`. + JavaScript-rendered, infinite-scroll or login-walled: a real browser driving the capture, which + in practice means `browsertrix-crawler` producing WACZ. Platform video: `yt-dlp`, which gets you + the metadata sidecar a screen recording never will. +4. **Hash on acquisition, before you look at the file.** A hash taken after you cropped, converted + or re-saved proves something about your working copy and nothing about the original. +5. **Timestamp the manifest, not every file.** One hash over the list of hashes, submitted to + OpenTimestamps or an RFC 3161 authority, binds the whole collection to a point in time at one + cost. That is the step that converts "I have a hash" into "I had this hash on that date". +6. **Write the provenance down in the same commit as the files.** URL, UTC time of capture, the + tool and version, who handed it to you, and how you reached the page. The route to a finding is + part of the finding, and it is the part you will forget first. +7. **Get a second independent capture.** Two archives of the same URL taken by different + infrastructure, or the same video from a second uploader, is the difference between a claim and + a corroborated claim. Two copies that both trace to one upload are one copy. + +Judgement calls worth naming: submitting a URL to a public archive creates a public record that +somebody is interested in it, so on a target that watches its referrers and its Wayback entries, +capture locally first and submit later. And a capture of a page you reached while logged in +contains your session — strip or redact before you share the WARC. + +## Key tools + +### wget (WARC capture) + +The standard way to get a byte-level record of an HTTP exchange rather than a rendered picture of +it. A WARC stores the request and response headers alongside the body for every resource fetched, +with a SHA-1 digest per record, so the file is a transcript of the conversation and not just its +outcome. Use it whenever the page is server-rendered; it is the most portable evidence format in +this area and every replay tool reads it. -## Public archives +```bash +brew install wget # or: apt install wget +``` ```bash -# push a URL into the Wayback Machine -curl -s "https://web.archive.org/save/https://example.com/page" +# the baseline capture: WARC transcript plus a CDX index of what went into it +wget --warc-file=capture-001 --warc-cdx \ + --page-requisites --adjust-extension --no-parent \ + 'https://example.com/page' + +# stamp the operator and case into the warcinfo record, so the file self-documents +wget --warc-file=capture-002 --warc-cdx \ + --warc-header="operator: J. Investigator" \ + --warc-header="description: case-2026-014, post cited in filing" \ + --warc-header="robots: off" \ + --page-requisites --adjust-extension 'https://example.com/page' + +# leave it uncompressed while you are inspecting it; grep works on a plain WARC +wget --warc-file=capture-003 --no-warc-compression 'https://example.com/page' + +# pull in the CDN that actually serves the images, or the capture renders blank +wget --warc-file=capture-004 --warc-cdx \ + --page-requisites --convert-links --adjust-extension \ + --span-hosts --domains example.com,cdn.example.com \ + 'https://example.com/page' + +# a whole section, politely: the delay is what stops you being blocked mid-capture +wget --warc-file=capture-005 --warc-cdx --recursive --level=2 --no-parent \ + --wait=2 --random-wait --limit-rate=500k \ + -U 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0 Safari/537.36' \ + 'https://example.com/newsroom/' + +# a list of URLs in one WARC, split at 1GB so the files stay movable +wget --warc-file=capture-006 --warc-max-size=1G --warc-cdx -i urls.txt -# check existing captures before assuming you need a new one -curl -s "http://archive.org/wayback/available?url=example.com/page&timestamp=20240101" | jq . +# second pass over the same site without re-storing what you already have +wget --warc-file=capture-007 --warc-dedup=capture-005.cdx --recursive --no-parent \ + 'https://example.com/newsroom/' + +# read back what the capture contains, without a replay tool +zcat capture-001.warc.gz | grep -a '^WARC-Target-URI:' | sort -u +awk 'NR>1 {print $5, $6, $1}' capture-001.cdx # status, SHA-1 digest, URL ``` -`archive.today` (also `archive.ph`, `archive.is`) renders JavaScript-heavy pages that Wayback often -fails on, and is markedly more resistant to takedown requests. Use both; they fail differently. +Read the CDX file: the digest column is a base32 SHA-1 of the response payload, which gives you a +free per-resource integrity check, and the status column tells you which requisites 404'd. A +capture whose CDX is full of `302` and `403` rows did not get the page. + +What this does not prove: the WARC records what *your* client received from *an* IP at that moment. +It does not prove the page looked that way to anyone else, does not survive geo-targeted or +personalised content, and carries no third-party attestation of the time — the `WARC-Date` is your +own clock. `--convert-links` rewrites the mirrored files on disk, not the WARC records, so the +transcript stays pristine while the browsable copy is usable; keep both. Above all, `wget` executes +no JavaScript, so on a modern single-page application it faithfully archives an empty shell. If the +page needs a browser, use one. + +### browsertrix-crawler -## Local capture +A headless Chrome driven by Webrecorder's crawler, writing WARC and WACZ. This is the live, +maintained answer for anything JavaScript-rendered, where `wget` returns a loading spinner. It +scrolls, autoplays, runs site-specific behaviours for the big platforms, and extracts page text for +search. + +The older Python crawlers in this niche have aged out: `wpull` carries a 2013–2016 copyright and no +release since 2019, and `grab-site` still pins Python 3.7/3.8 and was dropped from nixpkgs after +23.05. Neither is a reasonable thing to stand behind in 2026. Treat them as read-only history and +use the crawler below. ```bash -# single page with everything needed to render it offline -wget --page-requisites --convert-links --adjust-extension \ - --no-parent --span-hosts --domains example.com,cdn.example.com \ - 'https://example.com/page' +docker pull webrecorder/browsertrix-crawler +``` + +```bash +# single page, as WACZ, with text extracted for full-text search +docker run -v "$PWD/crawls:/crawls/" -it webrecorder/browsertrix-crawler crawl \ + --url 'https://example.com/post/123' --scopeType page \ + --generateWACZ --text to-pages --collection case-2026-014 + +# the page plus whatever it links to on the same host, two levels deep +docker run -v "$PWD/crawls:/crawls/" -it webrecorder/browsertrix-crawler crawl \ + --url 'https://example.com/newsroom/' --scopeType host --depth 2 \ + --generateWACZ --collection newsroom + +# a feed that only loads on scroll: give the behaviours time to finish +docker run -v "$PWD/crawls:/crawls/" -it webrecorder/browsertrix-crawler crawl \ + --url 'https://example.com/feed' --scopeType page-spa \ + --behaviors autoscroll,autoplay,autofetch,siteSpecific \ + --behaviorTimeout 300 --generateWACZ --collection feed + +# several URLs from a file, one collection, so the manifest covers the set +docker run -v "$PWD/crawls:/crawls/" -it webrecorder/browsertrix-crawler crawl \ + --urlFile /crawls/urls.txt --scopeType page --generateWACZ --collection batch-01 + +# the WACZ is a zip: look inside before you trust it +unzip -l crawls/collections/case-2026-014/case-2026-014.wacz +unzip -p crawls/collections/case-2026-014/case-2026-014.wacz datapackage-digest.json | jq . + +# the crawl log is JSON Lines — this is where blocked and failed pages surface +jq -r 'select(.logLevel=="error") | [.timestamp,.message] | @tsv' \ + crawls/collections/case-2026-014/logs/*.log +``` + +A WACZ is a WARC plus an index, a page list and a signed digest manifest, and it opens in +[ReplayWeb.page](https://replayweb.page/) with no server. That makes it the format to hand to +someone who is not going to install anything. Check `datapackage-digest.json` after every crawl: +if the crawler was blocked, you get a valid WACZ containing a block page, and nothing about the +file itself says so. + +What it does not tell you: a browser capture is a recording of one rendering session, which means +it includes whatever was personalised, A/B-tested or geo-fenced for that session. Running the same +crawl from a different exit and getting different content is a finding about the site, not an error. +Logged-in crawls bake the session into the archive — handle accordingly. + +### yt-dlp + +Platform video and audio with its metadata intact. The point is not the video file; a screen +recording gets you that. The point is the `.info.json` sidecar, which carries uploader ID, upload +date, duration, view and comment counts, available formats and often the original title and +description before anybody edited them. + +```bash +brew install yt-dlp # or: pipx install yt-dlp +``` + +```bash +# the provenance capture: video, metadata, description, thumbnail, subtitles +yt-dlp --write-info-json --write-description --write-thumbnail \ + --write-subs --sub-langs 'all' --no-mtime \ + -o '%(upload_date)s_%(uploader_id)s_%(id)s.%(ext)s' URL + +# what formats exist, before you commit to one +yt-dlp -F URL + +# the best single pre-merged file, so you archive bytes the platform served +yt-dlp -f 'best[ext=mp4]/best' --no-mtime URL + +# metadata only — enough to date and attribute a clip without downloading it +yt-dlp --skip-download --write-info-json --write-thumbnail URL + +# comments too; they routinely date an event more precisely than the upload does +yt-dlp --write-info-json --write-comments --skip-download URL + +# a whole channel, newest first, with an archive file so re-runs are incremental +yt-dlp --write-info-json --no-mtime --download-archive seen.txt \ + -o '%(upload_date)s_%(id)s.%(ext)s' 'https://example.com/@account/videos' + +# the intermediate pages the extractor fetched — useful when an extractor misreads a page +yt-dlp --write-pages --skip-download URL + +# the fields you will actually quote, straight out of the sidecar +jq -r '[.id,.uploader_id,.upload_date,.timestamp,.duration,.view_count,.title] | @tsv' *.info.json +``` + +`--no-mtime` matters: without it the file's mtime is set from the upload date, which silently +overwrites your own acquisition timeline with the platform's claim. `timestamp` in the sidecar is +Unix epoch seconds and is usually more precise than `upload_date`, which is date-only and rendered +in the platform's timezone of choice. + +What the sidecar does not establish: every field in it is the platform's assertion, repeated. +`upload_date` is when it was uploaded, never when it was filmed, and a re-upload resets it. +Extractors break when platforms change, so a field that is `null` may mean absent or may mean the +extractor lost it — check against the page. Rate limits and age or region gates will silently +truncate a channel pull; compare the count you got against the count the channel claims. + +### Wayback Machine, from the command line + +Three separate endpoints, used at three different moments. Availability answers "is there already a +capture"; CDX answers "what captures exist, and did the content change between them"; Save Page Now +creates a new one. + +```bash +# is there a capture at all, and what is the nearest one to a date +curl -s 'https://archive.org/wayback/available?url=example.com/page&timestamp=20240101' | jq . + +# every capture of a URL, with the payload digest that reveals real edits +curl -s 'https://web.archive.org/cdx/search/cdx?url=example.com/page&output=json&fl=timestamp,original,statuscode,digest,length' | jq -r '.[] | @tsv' + +# collapse consecutive identical captures: what is left is the change history +curl -s 'https://web.archive.org/cdx/search/cdx?url=example.com/page&output=json&fl=timestamp,digest&collapse=digest' | jq -r '.[] | @tsv' + +# everything ever captured under a host, which is how you find deleted pages +curl -s 'https://web.archive.org/cdx/search/cdx?url=example.com&matchType=domain&output=json&fl=timestamp,original,statuscode&limit=500' | jq -r '.[] | @tsv' + +# only the captures in a window, for a page that mattered on one specific day +curl -s 'https://web.archive.org/cdx/search/cdx?url=example.com/page&from=20260301&to=20260331&output=json&fl=timestamp,digest' | jq -r '.[] | @tsv' + +# submit a new capture: SPN2, authenticated with Internet Archive S3-style keys +curl -s -X POST -H 'Accept: application/json' \ + -H "Authorization: LOW $IA_ACCESS_KEY:$IA_SECRET_KEY" \ + -d 'url=https://example.com/page' -d 'capture_all=1' -d 'capture_screenshot=1' \ + 'https://web.archive.org/save' | jq . + +# poll the job until it reports success, and keep the returned timestamp +curl -s -H "Authorization: LOW $IA_ACCESS_KEY:$IA_SECRET_KEY" \ + "https://web.archive.org/save/status/$JOB_ID" | jq '{status, timestamp, original_url}' + +# fetch the archived copy itself, and hash what you got +curl -sL 'https://web.archive.org/web/20260301120000id_/https://example.com/page' | shasum -a 256 +``` + +Get the key pair from [archive.org/account/s3.php](https://archive.org/account/s3.php) and export +it as `IA_ACCESS_KEY` / `IA_SECRET_KEY`. **The unauthenticated `GET /save/<url>` form is no longer +usable for work** — an anonymous request to it returns HTTP 429 rather than a capture. Scripts that +still rely on it fail silently, because a 429 body looks like a page. + +The `id_` infix in that last URL asks for the original bytes without Wayback's navigation banner +injected, which is what you want if you intend to hash or diff the archived copy. The `digest` +column in CDX output is the payload hash: two captures sharing a digest are byte-identical, so +collapsing on it turns a thousand snapshots into the handful of moments the page actually changed. +That is the single most useful thing the CDX API does. + +What Wayback does not give you: permanence. A site can retroactively exclude itself, and captures +do disappear. It also honours robots directives at crawl time, so absence of a capture is not +evidence the page did not exist. And a Save Page Now capture is still just a capture — it attests +that the Internet Archive saw that content at that time, which is a strong third-party claim about +time and a weak one about authenticity. + +### archive.today + +A separate archive with a separate failure mode, which is exactly why you use it alongside Wayback. +It renders pages that Wayback flattens, ignores robots directives, and has a long record of not +removing things. There is no API and the submission form sits behind bot protection, so this one is +a browser job — do not script it. + +```text +1. Open https://archive.today (archive.ph and archive.is are the same service; + if one mirror is blocked where you are, try another) +2. Paste the URL into the lower box, "My url is alive and I want to archive its content" +3. Solve the challenge if you get one. Do not automate this step — it is the step + that gets the service to block your address. +4. Wait for the capture. The result URL looks like https://archive.ph/AbC12 + and the page header shows the capture time in UTC. +5. Click "screenshot" in the header to get the full-page render as a separate + artefact, and save it. The HTML capture and the screenshot fail differently. +6. Record the short URL AND the capture time. The short URL is your citation. +7. Before submitting, use the upper box to search for existing captures of the + same URL — the same reason you query Wayback's CDX API first. +``` + +Two things to know. The service is deliberately opaque about its infrastructure and funding, which +means you should not treat it as your only copy of anything; keep the local WARC. And it fetches +the page itself, from its own address, so a submission does not leak your IP to the target — but it +does create a public, searchable record that someone archived that URL at that minute. + +### Hashing and the manifest + +Hashes are the whole evidentiary argument. A hash recorded at acquisition and published or +timestamped separately lets you show, later, that the file you are producing is the file you took. +The manifest pattern — one file listing every hash, then one hash of the manifest — scales that to a +collection without timestamping a thousand files. + +```bash +# macOS ships shasum; coreutils ships sha256sum. Same output format. +shasum -a 256 evidence.mp4 +sha256sum evidence.mp4 +``` + +```bash +# hash the whole capture tree in a stable order, so the manifest is reproducible +find ./capture -type f -print0 | sort -z | xargs -0 shasum -a 256 > MANIFEST.sha256 + +# the manifest's own hash: this single line is what you timestamp and publish +shasum -a 256 MANIFEST.sha256 | tee MANIFEST.sha256.sha256 + +# verify the collection, any time later +shasum -a 256 -c MANIFEST.sha256 + +# show only what failed, which is what you actually want on a large set +shasum -a 256 -c MANIFEST.sha256 2>/dev/null | grep -v ': OK$' + +# a provenance record that travels with the files, in the same directory +cat > PROVENANCE.txt <<'EOF' +case: 2026-014 +url: https://example.com/post/123 +captured_at: 2026-10-04T04:07:56Z # UTC, from `date -u +%FT%TZ` +captured_by: J. Investigator +tooling: wget 1.25.0 --warc-file; yt-dlp 2026.08.19 +route: linked from https://example.com/newsroom/ on 2026-10-03 +notes: page required no login; no personalisation observed +EOF + +# fold the provenance into the manifest so the two cannot drift apart +shasum -a 256 PROVENANCE.txt >> MANIFEST.sha256 + +# find duplicates across two collections by hash, not by filename +sort MANIFEST.sha256 other/MANIFEST.sha256 | awk '{print $1}' | sort | uniq -d +``` + +Use SHA-256. MD5 and SHA-1 still appear in forensic tooling and are fine as identifiers, but a +collision-prone hash invites an argument you do not need to have. Note that `find | sort` matters: +an unsorted manifest changes order between runs and filesystems, which makes two manifests of the +same files look different. + +What a hash proves and does not: it proves the bytes have not changed since the hash was taken. It +says nothing about when they were taken, nothing about where they came from, and nothing about +whether the content is true. On its own a hash in your own notes is worth very little, because you +could have written it at any time — which is what the next section fixes. + +### Timestamping: OpenTimestamps and RFC 3161 + +The step that turns a hash into evidence. A timestamp is a third party's signed assertion that a +given digest existed before a given moment. Without it, your hash only proves internal +consistency; with it, you can show you held that exact file on that date and could not have +produced it later. + +```bash +pipx install opentimestamps-client # the `ots` command +# openssl is already present for the RFC 3161 route +``` + +```bash +# OpenTimestamps: anchors the digest in the Bitcoin blockchain, free, no account +ots stamp MANIFEST.sha256 + +# what the proof currently contains: pending calendar attestations, then a block +ots info MANIFEST.sha256.ots + +# calendars need an hour or so to get into a block; upgrade the proof afterwards +ots upgrade MANIFEST.sha256.ots + +# verify. Full verification wants a local Bitcoin node; without one you are +# trusting the calendar, which is weaker but still a third party +ots verify MANIFEST.sha256.ots -# video and audio, with metadata and subtitles preserved -yt-dlp --write-info-json --write-subs --write-thumbnail \ - --no-mtime -o '%(upload_date)s-%(id)s.%(ext)s' URL +# verify a digest you were given, without holding the file +ots verify -d 355c357708e4840d12b0d4284d9fe1a911c79c013953874f03068b81127f3363 MANIFEST.sha256.ots -# hash everything on acquisition -find ./capture -type f -exec sha256sum {} \; | tee capture.sha256 +# RFC 3161 route: build a timestamp query over the file's SHA-256 +openssl ts -query -data MANIFEST.sha256 -sha256 -cert -out manifest.tsq -# verify later -sha256sum -c capture.sha256 +# send it to a Time Stamping Authority and keep the signed reply +curl -s -H 'Content-Type: application/timestamp-query' \ + --data-binary @manifest.tsq https://freetsa.org/tsr -o manifest.tsr + +# read the token: the TSA's time, its identity and the digest it signed over +openssl ts -reply -in manifest.tsr -text + +# verify the token against the TSA's chain — this is the check a reviewer repeats +curl -sO https://freetsa.org/files/cacert.pem +curl -sO https://freetsa.org/files/tsa.crt +openssl ts -verify -data MANIFEST.sha256 -in manifest.tsr \ + -CAfile cacert.pem -untrusted tsa.crt +``` + +A successful verification prints `Verification: OK`, and that line is the thing you cite. Keep the +`.tsr`, the CA chain you verified against, and the exact file — the token is signed over the digest, +so a reviewer needs the file to recompute it. + +Which route to use: OpenTimestamps costs nothing, needs no account, and anchors to a public +blockchain, so the proof outlives any single company — but it is coarse (block granularity, roughly +an hour) and full verification wants a Bitcoin node. An RFC 3161 TSA gives you a precise, signed, +PKI-rooted token that institutional reviewers already recognise, at the price of trusting that TSA +and its certificate remaining validatable. Do both; they are cheap and they fail differently. + +What a timestamp does not prove: that the content is authentic, that the capture was complete, or +that you obtained it lawfully. It proves the file existed no later than that moment. That is a +narrow claim, and stating it narrowly is what makes it credible. + +### Bellingcat Auto Archiver + +The pipeline to set up when you are collecting continuously rather than case by case. It reads URLs +from a feeder, runs a platform-specific extractor, then runs enrichers that hash, thumbnail, +timestamp and push to Wayback, and writes the results to a database you can read as a case index. +It is the only tool here that performs the hash-and-timestamp discipline on every item without you +remembering to. + +```bash +pipx install auto-archiver +auto-archiver --help # the flag list is generated from the installed modules +``` + +Configuration is a single YAML file; the default path is `secrets/orchestration.yaml`. Only modules +named under `steps` are loaded, so the file is both configuration and a declaration of the pipeline. + +```yaml +# orchestration.yaml +steps: + feeders: [cli_feeder] + extractors: [generic_extractor] + enrichers: [hash_enricher, meta_enricher, thumbnail_enricher, + timestamping_enricher, opentimestamps_enricher] + databases: [csv_db, console_db] + storages: [local_storage] + formatters: [html_formatter] + +hash_enricher: + algorithm: SHA-256 # the only alternative is SHA3-512 + +timestamping_enricher: + tsa_urls: + - http://timestamp.identrust.com + - http://zeitstempel.dfn.de + allow_selfsigned: false # leaving this false is the whole point of the module + +local_storage: + save_to: ./archived + path_generator: url # flat | url | random + filename_generator: static # names the file after its hash + +csv_db: + csv_file: ./archived/results.csv + +logging: + level: INFO + file: ./archived/archiver.log +``` + +```bash +# one URL through the whole pipeline +auto-archiver --config orchestration.yaml 'https://example.com/post/123' + +# several at once; the CSV gets one row per URL +auto-archiver --config orchestration.yaml 'https://example.com/a' 'https://example.com/b' + +# no config file at all, for a one-off capture with sane defaults +auto-archiver --mode simple 'https://example.com/post/123' + +# override the enricher list without editing the file +auto-archiver --config orchestration.yaml \ + --enrichers hash_enricher timestamping_enricher opentimestamps_enricher \ + 'https://example.com/post/123' + +# push to Wayback as part of the run, with the same Internet Archive key pair +auto-archiver --config orchestration.yaml \ + --enrichers hash_enricher wayback_extractor_enricher \ + --wayback_extractor_enricher.key "$IA_ACCESS_KEY" \ + --wayback_extractor_enricher.secret "$IA_SECRET_KEY" \ + 'https://example.com/post/123' + +# feed a column of URLs from a spreadsheet export +auto-archiver --config orchestration.yaml \ + --feeders csv_feeder --csv_feeder.files urls.csv --csv_feeder.column link + +# turn up logging when an extractor silently returns nothing +auto-archiver --config orchestration.yaml --logging.level DEBUG 'https://example.com/post/123' + +# the results CSV is the case index: one row per URL, with hashes and archive links +column -s, -t < ./archived/results.csv | less -S +``` + +Two enrichers do the evidentiary work and they are not interchangeable. `timestamping_enricher` +aggregates the item's file hashes into one text file and gets an RFC 3161 token over it from each +TSA in the list; `opentimestamps_enricher` anchors the same digest via the OpenTimestamps +calendars. Note that `freetsa.org` is deliberately absent from the shipped TSA defaults because its +certificate does not validate against system authorities — you can add it, but only together with +`allow_selfsigned: true`, which weakens exactly the property you wanted. + +The `gsheet_feeder_db` module is what makes this a team tool: analysts paste URLs into a Google +Sheet, the archiver writes the hash, the stored path and the archive links back into adjacent +columns. The `wacz_extractor_enricher` shells out to Docker for browser-based capture, and +`wayback_extractor_enricher` needs the same key pair as Save Page Now. + +Watch the version. The client checks on startup and says so when it is behind, and per-platform +extractors break often enough that an old release produces empty captures that look like +successes. After every run, scan the results CSV for rows that carry a hash but no media. + +### ExifTool, for capture-time metadata + +Covered in depth on [Image & Video Forensics](/sheets/osint/image-video-forensics); the archiving +use is narrower. You are not asking whether the file is manipulated, you are recording what the +file asserted about itself at the moment it entered your custody, so that a later disagreement is +about the metadata rather than about your handling. + +```bash +# the full tag dump, archived next to the hash and never edited again +exiftool -j -g1 -a -u capture/evidence.mp4 > capture/evidence.mp4.exif.json + +# the fields that date and attribute a capture, in one line +exiftool -CreateDate -ModifyDate -FileModifyDate -GPSPosition -Make -Model -Software \ + capture/evidence.mp4 + +# every timestamp in the file, grouped by where it came from +exiftool -time:all -a -G1 -s capture/evidence.mp4 + +# the whole capture tree as one CSV, which doubles as a collection inventory +exiftool -r -csv -FileName -FileSize -MIMEType -CreateDate -GPSPosition ./capture/ > inventory.csv + +# flag files whose structure does not match their declared format +exiftool -r -validate -warning -a ./capture/ + +# strip metadata from a working copy before you publish, to protect a source +cp capture/evidence.jpg publish/evidence.jpg +exiftool -all= -overwrite_original publish/evidence.jpg +shasum -a 256 publish/evidence.jpg # a different file, so a different hash: say so ``` -`--no-mtime` matters: without it `yt-dlp` sets the file's modification time from the upload date, -which muddles your own acquisition timeline. +Run this on the copy, after hashing the original, and never with `-overwrite_original` on anything +in the capture tree. Note the last block: a published, stripped file is a *different artefact* with a +different hash, and conflating the two is a straightforward way to be accused of altering evidence. +Record both hashes and the relationship between them. -## Purpose-built tooling +What it does not tell you: metadata is written as easily as it is read, so a helpful +`CreateDate` is consistent-with and never proof. Most files pulled from platforms have had their +metadata stripped on upload, and that absence is normal rather than suspicious. + +## Case-management and pipeline tools | Tool | What it does | | --- | --- | @@ -83,14 +574,6 @@ which muddles your own acquisition timeline. | [Lumen](https://lumendatabase.org/) | Archive of takedown notices — sometimes the only record that something existed. | | [Web Archives](https://github.com/dessant/web-archives) | Browser extension that queries many archive services at once. | -Auto Archiver is the one to set up if you are collecting on an ongoing basis rather than -case-by-case: - -```bash -pipx install auto-archiver -auto-archiver --config orchestration.yaml -``` - ## Tool reference | Tool | What it does | Cost | @@ -111,6 +594,119 @@ auto-archiver --config orchestration.yaml nothing about the original. - **Archiving can notify.** Submitting a URL to a public archive creates a public record that someone is interested in it. +- **A capture that succeeded is not a capture that worked.** A WARC full of 403s, a WACZ containing + a block page, and a `yt-dlp` channel pull truncated by rate limiting all exit zero. Read the CDX + status column, the crawl log and the row count before you file anything. +- **Old scripts hitting `GET /save/<url>` now get HTTP 429.** Anonymous Save Page Now is no longer + usable; a 429 body is still a body, so a script that does not check the status silently records a + failure as a success. Use the authenticated SPN2 POST. +- **A hash in your own notes is nearly worthless.** You could have written it at any time. The + third-party timestamp is what makes it an assertion about the past rather than about your memory. + +## Worked example + +One datum: a URL to a post carrying a video clip, `https://example.com/@account/post/123`, cited in +a filing you have been asked to check. The URL, the hashes and the timestamps below are invented; +the order of the steps and what each one proves are not. + +```bash +# 1. what already exists, before you touch the page +curl -s 'https://web.archive.org/cdx/search/cdx?url=example.com/@account/post/123&output=json&fl=timestamp,digest,statuscode&collapse=digest' | jq -r '.[] | @tsv' +# timestamp digest statuscode +# 20260228183044 H4XQ... 200 +# 20260302094112 ZK7B... 200 +``` + +Two distinct payload digests, two days apart. Collapsing on digest is what revealed that: the page +was edited between 28 February and 2 March, which is a finding in itself and the reason to pull +both old captures before capturing the live page. + +```bash +# 2. local capture of the live page, byte-level, with the case in the warcinfo record +mkdir -p case-2026-014 && cd case-2026-014 +wget --warc-file=live --warc-cdx --no-warc-compression \ + --warc-header="operator: J. Investigator" \ + --warc-header="description: case-2026-014, post cited at para 17" \ + --page-requisites --adjust-extension --no-parent \ + 'https://example.com/@account/post/123' +awk 'NR>1 {print $5, $1}' live.cdx | sort | uniq -c +# 14 200 https://example.com/... +# 3 403 https://cdn.example.com/media/... +``` + +Three requisites returned 403, and they are the media files. The `wget` capture has the page text +and not the clip, which is the common outcome and the reason the next two steps exist rather than +being optional extras. + +```bash +# 3. the rendered page, in a browser, because the post body loads via JavaScript +docker run -v "$PWD/crawls:/crawls/" -it webrecorder/browsertrix-crawler crawl \ + --url 'https://example.com/@account/post/123' --scopeType page \ + --generateWACZ --text to-pages --collection case-2026-014 +unzip -p crawls/collections/case-2026-014/case-2026-014.wacz datapackage-digest.json | jq . +# { "path": "datapackage.json", "hash": "sha256:9f2c...", "signedData": null } +``` + +`signedData: null` means the WACZ is unsigned — fine, because you are about to timestamp it +yourself. The digest over `datapackage.json` is the crawler's own integrity check on the archive; +record it, then stop relying on it and use your own hash. + +```bash +# 4. the media, with its provenance sidecar, which the WARC never got +yt-dlp --write-info-json --write-description --write-thumbnail --no-mtime \ + -o '%(upload_date)s_%(uploader_id)s_%(id)s.%(ext)s' \ + 'https://example.com/@account/post/123' +jq -r '[.id,.uploader_id,.upload_date,.timestamp,.duration,.view_count] | @tsv' *.info.json +# 123 account 20260226 1772120000 47 18422 +``` + +The sidecar says the clip was uploaded on 26 February — two days *before* the earliest Wayback +capture of the post, and before the edit the CDX diff exposed. That is the pivot: the video predates +the post text it now sits under, so the claim to check is whether the caption was changed, not +whether the footage is real. + +```bash +# 5. hash everything, in a stable order, before any further handling +cd .. && find ./case-2026-014 -type f -print0 | sort -z | xargs -0 shasum -a 256 > MANIFEST.sha256 +wc -l MANIFEST.sha256 && shasum -a 256 MANIFEST.sha256 +# 31 MANIFEST.sha256 +# 4b1c8e... MANIFEST.sha256 +``` + +```bash +# 6. the step that makes the hash mean something: two independent timestamps +ots stamp MANIFEST.sha256 +openssl ts -query -data MANIFEST.sha256 -sha256 -cert -out manifest.tsq +curl -s -H 'Content-Type: application/timestamp-query' \ + --data-binary @manifest.tsq https://freetsa.org/tsr -o manifest.tsr +openssl ts -reply -in manifest.tsr -text | grep -E 'Status:|Time stamp:' +# Status: Granted. +# Time stamp: Oct 4 04:06:07 2026 GMT +``` + +Now the claim is narrow and defensible: this set of 31 files, with these hashes, existed no later +than 04:06 UTC on 4 October 2026, attested by a TSA and by the Bitcoin blockchain, neither of which +you control. + +```bash +# 7. a third-party capture, submitted last so the local copy exists first +curl -s -X POST -H 'Accept: application/json' \ + -H "Authorization: LOW $IA_ACCESS_KEY:$IA_SECRET_KEY" \ + -d 'url=https://example.com/@account/post/123' -d 'capture_all=1' \ + 'https://web.archive.org/save' | jq -r '.job_id' +# spn2-9f3c... +``` + +Then the same URL by hand into [archive.today](https://archive.today), and its short URL and UTC +capture time into `PROVENANCE.txt`. Submitting last is deliberate: both submissions are public, so +if the account notices and deletes the post, you already hold the capture. + +What this run establishes: the page as it stood at a timestamped moment, the clip with the +platform's own metadata, and a change history from Wayback showing the post was edited after the +video was uploaded. What it does not establish: that the footage shows what the caption claims, or +where or when it was filmed. Those are +[Image & Video Forensics](/sheets/osint/image-video-forensics) and +[Geolocation](/sheets/osint/geolocation) questions, and no amount of archiving substitutes for them. ## Broader catalogues diff --git a/src/content/sheets/osint/data-analysis-and-visualisation.md b/src/content/sheets/osint/data-analysis-and-visualisation.md @@ -4,9 +4,9 @@ description: "Clean messy collected data, map relationships between entities, bu category: osint subcategory: "Archiving & Analysis" tags: [osint, analysis, visualisation, graphs, timelines] -tools: [openrefine, gephi, datasette, datawrapper, qgis] +tools: [openrefine, csvkit, sqlite-utils, datasette, pandas, networkx, gephi, datawrapper, pinpoint] difficulty: intermediate -updated: 2026-09-28 +updated: 2026-10-04 references: - name: "Bellingcat's Online Investigation Toolkit" url: "https://bellingcat.gitbook.io/toolkit" @@ -24,81 +24,471 @@ references: ## What this covers -The part after collection. Thousands of rows of scraped posts, a list of company officers, a set of -geotagged images — none of it means anything until it is cleaned, related and shown. Analysis is -also where most errors enter an investigation, because a chart makes a weak claim look strong. +The part after collection, where a pile of scraped rows becomes a claim you are willing to publish. +Cleaning, relating, sequencing and showing are four separate jobs with four different failure +modes, and analysis is where most errors enter an investigation because a chart makes a weak claim +look strong. Every step below either loses rows or changes their meaning, so the discipline is to +know which, and to write it down. + +## Method + +1. **Count the rows you started with, and count them again after every step.** A step that drops + rows silently is the single most common way a dataset ends up saying the wrong thing. Record the + number before and after; if it changed and you did not intend it, stop. +2. **Normalise before you deduplicate, deduplicate before you count.** `Northgate Ltd`, + `Northgate Limited` and `northgate ltd` are three rows and one company. Any count you compute + before merging them is wrong, and no later step fixes it. +3. **Disable type inference on first contact.** Tools guess, and a date guessed month-first turns + 1 March into 3 January without a warning. Load everything as text, inspect, then convert + deliberately with a format you specified. +4. **Decide what the unit of analysis is before you join.** One row per post, per account, per + account-day? A join that silently fans out one-to-many inflates every count downstream, and the + resulting chart looks fine. +5. **Keep raw, cleaned and derived as three separate files.** The raw file is never edited. The + cleaning is a script, not a sequence of clicks you cannot repeat. The derived file is what the + chart reads. +6. **Mark the gaps.** A collection outage renders as a quiet period and reads as an event. If you + know when your collection was incomplete, that belongs on the timeline as a shaded band, not as + a footnote nobody reads. +7. **Choose the figure that makes the weakest honest claim.** If the finding is "these two accounts + post within sixty seconds of each other, 94 times", say that. A force-directed graph of the whole + dataset with those two nodes somewhere in it is prettier and claims more than you can support. + +## Key tools + +### OpenRefine + +The cleaning tool to learn properly, because its clustering is the one feature that no CLI +replicates well: it proposes groups of values that are probably the same thing and lets you merge +each group in one click, with the count of affected rows visible as you decide. That visibility is +the point — merging is a judgement, and OpenRefine makes you make it row-count in hand. + +Download the current release (3.10.0, February 2026) from +[openrefine.org/download](https://openrefine.org/download) and run it locally; it serves a browser +UI on `127.0.0.1:3333` and your data never leaves the machine. + +Clustering lives under a column's dropdown, `Edit cells` → `Cluster and edit`. Work the methods in +this order: + +```text +key collision / fingerprint lowercases, strips punctuation and diacritics, sorts + the words. Start here: almost no false positives. + Catches "SMITH, John" = "John Smith". +key collision / ngram-fingerprint n=2 catches typos and missing spaces that fingerprint + misses. Raise n to loosen, lower it to tighten. +key collision / metaphone3 English phonetics: "Stephen" = "Steven". Use + cologne-phonetic for German, daitch-mokotoff for + Slavic and Yiddish names, beider-morse for a stricter pass. +nearest neighbour / levenshtein edit distance, with a radius you set. Slow. This is + where real false positives start, so read every group. +nearest neighbour / ppm compression-based, for long strings. Last resort; it + over-merges short values badly. +``` + +GREL, in the `Edit cells` → `Transform` box, is for everything clustering cannot do: + +```text +value.trim() strip the whitespace that breaks every join +value.replace(/\s+/, " ").trim() collapse runs of whitespace to one space +fingerprint(value) the clustering key itself, as a new column, + so you can group on it in SQL later +value.replace(/\b(Ltd|Limited|PLC|LLC|Inc)\b\.?/i, "").trim() + drop corporate suffixes before matching names +value.toDate("dd/MM/yyyy") parse an unambiguous day-first date explicitly +value.toDate(false, "dd/MM/yyyy", "yyyy-MM-dd") try day-first, then ISO; false = not month-first +value.toDate("dd/MM/yyyy").datePart("hours") pull a component back out +diff(value.toDate("yyyy-MM-dd"), cells["first_seen"].value.toDate("yyyy-MM-dd"), "days") + day gap between two columns +value.match(/^(\+?\d{1,3})[\s-]?(\d{6,})$/)[1] capture group from a regex, as a value +value.splitByCharType().join("|") expose where letters and digits meet, which + is how you find "acct12" vs "acct 12" +cross(value, "officers", "company_name")[0].cells["role"].value + look a value up in another OpenRefine project +``` + +Reconciliation is the other thing worth the setup: it matches a column of names against an external +authority and returns scored candidates rather than a guess. Add a service under +`Reconcile` → `Start reconciling` → `Add standard service`: + +```text +https://wikidata.reconci.link/en/api Wikidata; swap "en" for another language code +``` + +That host now 307-redirects to `wikidata-reconciliation.wmcloud.org`, which is the current canonical +home — either URL works, the second avoids the hop. The older +`wdreconcile.toolforge.org` service is deprecated and unreliable; do not point new projects at it. + +Read reconciliation output as candidates, never as answers. The service returns a match score and +OpenRefine auto-matches above a threshold, which means it will confidently bind a small company to +a famous one of a similar name. Set the threshold high, review the auto-matches, and keep the +unmatched rows visible rather than dropping them — an unmatched name is information about your +authority's coverage, not about the entity. + +### csvkit + +The right tool between "I have a CSV" and "I have a database": it does the inspection, cutting, +filtering and joining that would otherwise be a throwaway script, and it streams, so it works on +files larger than memory. Reach for it before pandas when the job is shaped like a shell pipeline. + +```bash +pipx install csvkit +``` + +```bash +# what are the columns actually called, and in what order +csvcut -n collected.csv -## Cleaning +# does the file parse at all: ragged rows, empty columns, mismatched header +csvclean -a --label - collected.csv > /dev/null -Collected data is always dirty: inconsistent name spellings, mixed date formats, duplicate entities -under slightly different labels. +# a readable look at the first rows, before you decide anything +csvcut -c 1-6 collected.csv | csvlook --max-rows 20 -**OpenRefine** is the right tool and is underused. Its clustering function finds values that are -probably the same thing — `Jon Smith`, `John Smith`, `SMITH, John` — and lets you merge them in one -pass. Do this before any counting, or your counts are wrong. +# the whole statistical profile: type, nulls, unique count, min, max, frequencies +csvstat collected.csv + +# just the questions that matter on first contact +csvstat --nulls collected.csv +csvstat -c company --freq --freq-count 25 collected.csv + +# convert from whatever you were given, naming the sheet explicitly +in2csv -f xlsx --sheet 'Payments' register.xlsx > payments.csv +in2csv -f json -k results api_dump.json > api.csv + +# filter by regex, case-insensitively, keeping the header +csvgrep -c company -r '(?i)northgate' collected.csv > northgate.csv + +# rows where ANY column matches, for a name you cannot locate in a known column +csvgrep -a -r '(?i)northgate' collected.csv + +# join two collections on an id, keeping every left row so you can see the misses +csvjoin -I -c id --left payments.csv officers.csv > merged.csv + +# stack files from separate collection runs, tagging each with its source +csvstack -g 2026-02,2026-03 -n batch feb.csv mar.csv > all.csv + +# ad-hoc SQL over CSVs without creating a database +csvsql --query "SELECT company, COUNT(*) n, SUM(amount) total + FROM payments GROUP BY 1 HAVING n > 1 ORDER BY total DESC" payments.csv +``` + +**Use `-I` on first contact, every time.** csvkit infers types by default, and the inference is +month-first: a column containing `01/03/2026` comes out of a plain `csvjoin` as `2026-01-03`, so +1 March silently becomes 3 January, with no warning and no error. `-I` passes the value through +verbatim. If you do want parsing, say what the format is with `--date-format '%d/%m/%Y'` — and note +that a column with genuinely mixed formats will defeat that too, which is a reason to split the +column rather than to trust the parser. + +Two other limits worth knowing: `csvjoin` reads every input fully into memory and says so in its +own help, so a multi-gigabyte join belongs in SQLite rather than here; and the `--blanks` flag +exists because csvkit converts `""`, `na`, `n/a`, `none` and `.` to NULL by default, which is +usually right and is occasionally the destruction of a meaningful category. + +### sqlite-utils + +The step that turns a directory of CSVs into something you can actually interrogate. It builds the +schema from the data, handles the inserting, and gives you full-text search and foreign keys +without writing DDL. Past roughly a hundred thousand rows this is where analysis should live rather +than in a dataframe. + +```bash +pipx install sqlite-utils +``` + +```bash +# CSV straight into a table, with a primary key so re-runs are idempotent +sqlite-utils insert case.db payments payments.csv --csv --pk id + +# everything as TEXT, which is the safe default when the dates are ambiguous +sqlite-utils insert case.db payments_raw payments.csv --csv --no-detect-types + +# newline-delimited JSON, e.g. a scraper's output, flattening nested objects +sqlite-utils insert case.db posts posts.jsonl --nl --flatten --pk id + +# re-run a collection without duplicating: replace rows that already exist +sqlite-utils insert case.db posts new.jsonl --nl --pk id --replace + +# add a column from the second batch without rebuilding the table +sqlite-utils insert case.db posts batch2.jsonl --nl --pk id --alter + +# inspect what you built +sqlite-utils schema case.db +sqlite-utils tables case.db --counts + +# full-text search across the columns that carry names and free text +sqlite-utils enable-fts case.db payments name company --create-triggers +sqlite-utils search case.db payments 'northgate' -c id -c company + +# the normalisation that makes counting honest: one row per distinct company +sqlite-utils extract case.db payments company --table companies --fk-column company_id + +# a query, as JSON, straight out of the CLI +sqlite-utils case.db "SELECT company, COUNT(*) n FROM payments GROUP BY 1 ORDER BY n DESC LIMIT 20" + +# throwaway query against CSVs with no database at all +sqlite-utils memory payments.csv officers.csv \ + "SELECT p.company, o.role FROM payments p JOIN officers o ON p.id = o.id" + +# indexes, once a query starts being slow rather than instant +sqlite-utils create-index case.db payments company date +``` + +`--pk` is the one flag to never omit: without it every re-run of a collection appends duplicates, +and you will not notice until a count is wrong. Note that type detection is on by default — the +flag is `--no-detect-types` to turn it off, and there is no `--detect-types`. `extract` is the +underused one: it pulls a repeated text column into its own table with a foreign key, which is both +the normalisation step and a cheap way to see how many distinct entities you really have. + +What it does not do: validate anything. `sqlite-utils` will happily build a table where `amount` is +`TEXT` because one row contained `n/a`, and every `SUM` over it then returns zero without +complaining. Check the schema after every insert. + +### Datasette + +Publishing a dataset other people can query, without building an application. Point it at the +SQLite file and you get faceted browsing, arbitrary SQL in the URL, CSV and JSON export, and a +permalink per row — which means a colleague can cite a specific record and you can both see the +query that produced a number. + +```bash +pipx install datasette +``` + +```bash +# serve locally, read-only, and open a browser +datasette serve -i case.db -o + +# bind somewhere a colleague on the network can reach, with a port you chose +datasette serve -i case.db -h 0.0.0.0 -p 8080 + +# several databases at once, with cross-database joins enabled +datasette serve -i case.db -i reference.db --crossdb + +# source, licence and column descriptions, which is the difference between a +# dataset and a pile of rows +datasette serve -i case.db -m metadata.yml + +# ad-hoc: query CSVs in memory without building a database first +datasette --memory + +# a scripted query against the JSON API, for a figure you will regenerate +datasette serve -i case.db --get '/case.json?sql=select+company,count(*)+from+payments+group+by+1' + +# publish to a container host +datasette publish cloudrun case.db --service case-2026-014 +datasette publish heroku case.db --name case-2026-014 +``` -For anything scriptable, pandas: +`-i` opens the file immutable, which is what you want for published evidence: nothing served can +alter the database, and Datasette can cache aggressively because the contents cannot change. + +Note that `datasette publish` ships only `cloudrun` and `heroku` targets. Vercel, Fly and Cloud Run +with custom settings are plugins (`datasette-publish-vercel`, `datasette-publish-fly`) that must be +installed first — a `datasette publish vercel` invocation fails with an unknown-command error on a +bare install. Current stable is the 0.6x line; the 1.0 alphas change the plugin and metadata APIs, +so pin your version if you are publishing something you need to rebuild identically later. + +The caution with Datasette is social rather than technical: a published database is far more +exposing than a chart. Every row is readable, every join is possible, and arbitrary SQL means +anybody can compute something you did not anticipate. Before publishing, drop the columns you only +needed for cleaning, and remember that a `source_url` column can re-identify people a redacted name +column was meant to protect. + +### pandas, for timelines + +Sequence is the most load-bearing and most abused structure in open-source work: who posted first, +how fast something propagated, whether an account's activity pattern changed. Build it in pandas +because the operations you need — timezone conversion, resampling, gap measurement — are one line +each and are exactly the ones a spreadsheet gets wrong. + +```bash +pipx install pandas +``` ```python import pandas as pd -df = pd.read_csv("collected.csv") -df["date"] = pd.to_datetime(df["date"], errors="coerce", utc=True) -df["name"] = df["name"].str.strip().str.casefold() -df = df.drop_duplicates(subset=["name", "date"]) +df = pd.read_csv("collected.csv", dtype=str) # everything as text; convert deliberately -# what did parsing fail on? this is where silent data loss hides -print(df["date"].isna().sum(), "unparseable dates") -``` +# parse explicitly, and keep the failures visible rather than dropping them +df["ts"] = pd.to_datetime(df["ts"], errors="coerce", utc=True, format="mixed") +bad = df["ts"].isna() +print(f"{bad.sum()} of {len(df)} timestamps unparseable") +df.loc[bad, ["id", "ts"]].to_csv("unparsed_timestamps.csv", index=False) +df = df.dropna(subset=["ts"]).sort_values("ts") -Always check what failed to parse. Rows silently dropped by a coercion are the classic way a -dataset ends up telling you the wrong thing. +# volume per day, with empty days present as zero rather than missing +print(df.set_index("ts").resample("1D").size()) -## Relationships +# posting hours in the timezone the actor plausibly lives in, not in UTC +local = df["ts"].dt.tz_convert("Europe/Kyiv") +print(local.dt.hour.value_counts().sort_index()) -**Gephi** for network graphs — people, companies, accounts and the edges between them. The useful -outputs are usually degree (who is most connected), betweenness (who bridges otherwise separate -clusters) and modularity (what the communities are). Force-directed layouts look impressive and say -little on their own; the metrics are the finding. +# activity heatmap: hour of day by actor, which is how a shift pattern shows up +print(df.groupby([local.dt.hour, "actor"]).size().unstack(fill_value=0)) -**Maltego** automates collection *into* a graph via transforms, which is convenient but ties you to -its data sources. +# gaps between consecutive events, in seconds — the coordination signal +df["gap_s"] = df["ts"].diff().dt.total_seconds() +print(df.nsmallest(20, "gap_s").loc[:, ["ts", "actor", "gap_s"]]) -For a company-ownership chain, a graph is usually overkill — a simple parent/subsidiary tree is -clearer and harder to misread. +# near-simultaneous posts by different actors, which is the claim worth making +cols = ["ts", "actor", "id"] +pairs = pd.merge_asof( + df.loc[:, cols].rename(columns={"actor": "a", "id": "id_a"}), + df.loc[:, cols].rename(columns={"actor": "b", "id": "id_b"}), + on="ts", tolerance=pd.Timedelta("60s"), direction="nearest", allow_exact_matches=False, +) +print(pairs.query("a != b").groupby(["a", "b"]).size().sort_values(ascending=False).head(20)) -## Querying +# known collection gaps, recorded as data so they reach the chart +gaps = pd.DataFrame({"start": pd.to_datetime(["2026-03-05T00:00Z"]), + "end": pd.to_datetime(["2026-03-07T12:00Z"])}) +gaps.to_csv("collection_gaps.csv", index=False) -**Datasette** turns a SQLite file into a browsable, queryable, publishable web interface. It is the -fastest route from "I have a CSV" to "colleagues can explore this". +df.to_csv("timeline.csv", index=False) # the derived file the chart reads +``` + +**`format="mixed"` is not optional.** Without it, pandas infers a single format from the first +value and silently coerces every differently-formatted-but-perfectly-valid value to `NaT`: a column +holding `2026-03-01T08:14:00Z` and `2026-03-01 09:02` loses the second row entirely under +`errors="coerce"`. That is real data loss that reports as a clean parse, which is why the row count +and the `unparsed_timestamps.csv` dump above exist. + +Two more traps. `tz_convert` on a naive column raises; `tz_localize` first if the source had no +offset, and if you do not know the source timezone, say so rather than assuming UTC. And a +`resample` over a sparse series fabricates the zeros it displays, which is correct for volume and +wrong for anything you then average. + +### networkx + +Build and measure the graph in code, then hand it to Gephi to look at. Doing the metrics here +rather than in the GUI is what makes them reproducible: the numbers come out of a script you can +re-run and diff, instead of a sequence of panel clicks nobody recorded. ```bash -pipx install datasette sqlite-utils +pipx install networkx +``` -sqlite-utils insert data.db records collected.csv --csv -datasette data.db -datasette publish vercel data.db --project my-investigation +```python +import networkx as nx +import pandas as pd + +edges = pd.read_csv("edges.csv") # source, target, weight, kind +G = nx.from_pandas_edgelist(edges, "source", "target", + edge_attr=["weight", "kind"], create_using=nx.DiGraph) +print(G.number_of_nodes(), "nodes,", G.number_of_edges(), "edges") + +# who is most connected, in and out separately — they mean different things +print(sorted(G.in_degree(weight="weight"), key=lambda x: -x[1])[:10]) +print(sorted(G.out_degree(weight="weight"), key=lambda x: -x[1])[:10]) + +# who bridges otherwise separate clusters: usually the actual finding +btw = nx.betweenness_centrality(G, normalized=True) +print(sorted(btw.items(), key=lambda x: -x[1])[:10]) + +# communities, computed on the undirected view as modularity requires +communities = nx.community.greedy_modularity_communities(G.to_undirected()) +print([len(c) for c in communities]) + +# is the graph one thing or several disconnected pieces +print([len(c) for c in nx.weakly_connected_components(G)]) + +# the ownership or reply chain between two specific entities +print(nx.shortest_path(G, "acct_a", "acct_z")) + +# a subgraph two hops out from one node, which is what you actually draw +ego = nx.ego_graph(G, "acct_a", radius=2) +print(ego.number_of_nodes(), "nodes in the two-hop neighbourhood") + +# write metrics back onto the nodes, then export for Gephi +nx.set_node_attributes(G, btw, "betweenness") +nx.set_node_attributes(G, dict(G.degree()), "degree") +nx.write_gexf(G, "case.gexf") ``` -`sqlite-utils` alone is worth learning — it handles the CSV-to-database step that otherwise eats an -afternoon. +Compute the metrics on the graph you can defend, not the one you collected. Betweenness on a graph +built from "both accounts used the same hashtag" measures hashtag popularity and nothing else; +betweenness on a graph of declared company directorships measures something real. The edge +definition is the entire analysis, and it belongs in the caption. + +`betweenness_centrality` is roughly O(nodes × edges) and becomes impractical in the tens of +thousands of nodes — pass `k=500` to sample pivots and accept an approximation. Note that +`greedy_modularity_communities` needs the undirected view, and that community detection is +stochastic in most implementations: run it twice, and if the communities differ materially, the +structure is not there. + +### Gephi + +The place to look at the graph, after the metrics are computed. Current line is 0.11, free and +open source, and it reads the `.gexf` written above with node attributes intact. + +```text +1. File → Open → case.gexf. Check the import report for discarded edges. +2. Data Laboratory tab first, not Overview. Confirm the node and edge counts + match what networkx printed. A mismatch means parallel edges were merged. +3. Appearance → Nodes → Size → Ranking → betweenness (the attribute you wrote + from networkx). Size by the metric; never size by eye. +4. Appearance → Nodes → Colour → Partition → modularity_class if you computed + it in Gephi, or your own community column from networkx. +5. Layout → ForceAtlas2. Enable "Prevent Overlap" and "LinLog mode" only once + it has settled. Stop it; do not let it run to a shape you like. +6. Filters → Topology → Degree Range to drop the degree-1 fringe, which is + usually most of the nodes and none of the finding. +7. Preview tab for the export. Turn node labels on only for the nodes you will + name in the text. +``` -## Timelines and maps +The honest use of Gephi is to communicate a structure you already established numerically. A +force-directed layout is not a measurement: the distance between two nodes on screen has no units, +clusters appear in random data, and re-running the layout moves everything. If the figure is doing +the arguing, the figure is overclaiming — put the degree and betweenness numbers in the caption and +let the picture be an illustration. + +### Datawrapper + +Publishing a chart that survives being read on a phone by someone who does not trust you. It +handles the things hand-rolled charts get wrong — colour-blind-safe palettes, responsive layout, +accessible axis labelling, a visible source line — and the output is an embed plus a static +fallback. Free tier covers everything below; the paid tiers add custom themes and private +workspaces. + +```text +1. Upload: paste the derived CSV (timeline.csv, not collected.csv) into the + Upload Data step. Check the "Check & Describe" screen — it shows the parsed + type per column, and this is where a month-first date misread surfaces + before it reaches the chart. +2. Visualize → chart type. Lines for a rate over time, bars for counts by + category, a symbol map only if position is the finding. +3. Refine → Appearance: leave the default palette. Its colours are chosen for + colour-vision deficiency and for greyscale printing. +4. Refine → Axes: a bar chart's value axis starts at zero. Non-zero baselines + belong only on line charts, and then with the baseline labelled. +5. Annotate → Title, description, source and byline. The source field is not + optional: name the dataset, the date collected, and the number of records. +6. Annotate → highlight the specific data points you discuss in the text, and + add a shaded range for every known collection gap. +7. Publish & Embed → take both the responsive iframe and the static PNG. The + PNG is what survives in a PDF, an email and an archive of your own article. +``` -- **[Time.Graphics](https://time.graphics/)** — quick shareable timelines. -- **[Pinpoint](https://journaliststudio.google.com/pinpoint/)** — OCR and entity extraction across - large document sets; finds names and places across thousands of scanned pages. -- **QGIS** — for anything where position matters. See - [Maps & Satellite Imagery](/sheets/osint/maps-and-satellite-imagery). +The failure mode is not the tool, it is reaching for a chart too early. A figure built from four +observations is a figure about four observations, and "50% increase" on a base of four is noise +rendered at 300 dpi. State the denominator in the subtitle, every time. -## Publishing figures +### Maps and spatial data -**Datawrapper** produces clean, accessible, responsive charts and maps with almost no effort, and -handles the things hand-rolled charts get wrong — colour-blind-safe palettes, mobile layout, proper -axis labelling. **RAWGraphs** covers the less common chart types. +Anything where position is the finding belongs in QGIS and GDAL, which are covered properly on +[Maps & Satellite Imagery](/sheets/osint/maps-and-satellite-imagery) — including `qgis_process` +for scripted geoprocessing and the `gdal_translate` / `gdalwarp` workflow. Use that sheet rather +than a second copy here. The one rule that belongs on this page: a coordinate column is not spatial +data until you know its CRS, and measuring distances in degrees because the layer is in `EPSG:4326` +is the standard way to publish a wrong number. -Whatever you use: label the axes, state the source, show the sample size, and do not start a bar -chart's axis anywhere but zero. +For document sets rather than coordinates, +[Pinpoint](https://journaliststudio.google.com/pinpoint/) does OCR, transcription and entity +extraction across thousands of scanned pages, and is free for journalists and researchers on +application. It is the fastest route from a PDF dump to a searchable corpus, and its entity +extraction is a lead generator, not a finding. ## Tool reference @@ -129,6 +519,139 @@ chart's axis anywhere but zero. - **Pretty visualisations oversell weak data.** The more convincing the figure, the more carefully the caveats need stating. - **Small numbers do not support percentages.** "50% increase" on a base of four is noise. +- **Type inference is a silent rewrite.** csvkit reads `01/03/2026` month-first and emits + `2026-01-03`; pandas without `format="mixed"` coerces every value that does not match the first + row's format to `NaT`. Load as text, convert deliberately, and print the row count either side. +- **A null is not a zero.** Columns with missing values make `COUNT(*)` and `COUNT(col)` disagree, + and a mean computed over the non-nulls reported against the full row count understates by exactly + the share that was missing. + +## Worked example + +One datum: `collected.csv`, 12,400 rows scraped from a set of accounts, with columns +`id,ts,actor,company,amount,url`. The numbers below are invented; the order of the steps, and what +each one reveals or destroys, is not. + +```bash +# 1. before anything: does it parse, and how many rows are there really +wc -l collected.csv +csvcut -n collected.csv +csvclean -a --label - collected.csv > /dev/null +# 12401 collected.csv +# 1: id 2: ts 3: actor 4: company 5: amount 6: url +# 14 rows were longer than the header row +``` + +Fourteen ragged rows. Those are almost always unescaped commas inside a free-text field, which +means the data in them is shifted one column right — `amount` holding a URL fragment. Fix or +exclude them now, because every count below would otherwise be wrong by fourteen in an unknown +direction. + +```bash +# 2. the profile, with inference OFF so nothing is reinterpreted on the way in +csvstat --nulls collected.csv +csvstat -c company --freq --freq-count 15 collected.csv +# amount: True (1,902 nulls) +# { "Northgate Ltd": 412, "Northgate Limited": 198, "northgate ltd": 61, +# "NORTHGATE LTD.": 44, "Westvale PLC": 390, ... } +``` + +Four spellings of one company, 715 rows between them. Counted raw, Northgate is the fourth-largest +counterparty; counted merged, it is the largest. That single merge changes the headline, which is +exactly why clustering comes before counting. + +```text +# 3. OpenRefine: cluster the company column +Edit cells → Cluster and edit → key collision / fingerprint + → 1 cluster, 4 values, 715 rows → merge to "Northgate Ltd" +Then: key collision / ngram-fingerprint (n=2) + → 1 further cluster: "Westvale PLC" + "Westvale P.L.C" (390 + 7 rows) +Then: nearest neighbour / levenshtein, radius 2 + → proposes "Northgate Ltd" + "Northgate Mining Ltd" — REJECT. + Different companies. Export the cleaning operations as JSON so the + merge is reproducible and the rejection is on the record. +``` + +The levenshtein rejection is the part worth recording. An automatic merge at that radius would have +folded two real companies into one and inflated the headline figure by another 130 rows, and nothing +downstream would have shown it. + +```bash +# 4. into SQLite, as text first, so the ambiguous dates survive the trip +sqlite-utils insert case.db payments cleaned.csv --csv --pk id --no-detect-types +sqlite-utils schema case.db +sqlite-utils case.db "SELECT company, COUNT(*) n, COUNT(amount) with_amount + FROM payments GROUP BY 1 ORDER BY n DESC LIMIT 5" +# [{"company": "Northgate Ltd", "n": 715, "with_amount": 601}, +# {"company": "Westvale PLC", "n": 397, "with_amount": 397}, ...] +``` + +114 Northgate rows have no amount. If you had summed without checking, the mean would have been +computed over 601 rows and reported over 715 — a 19% understatement presented as a fact. + +```python +# 5. the timeline, with the parse failures written out rather than dropped +import pandas as pd +df = pd.read_csv("cleaned.csv", dtype=str) +df["ts"] = pd.to_datetime(df["ts"], errors="coerce", utc=True, format="mixed") +print(df["ts"].isna().sum(), "of", len(df), "unparseable") # 38 of 12386 +df.loc[df["ts"].isna(), ["id", "ts", "url"]].to_csv("unparsed_timestamps.csv", index=False) +df = df.dropna(subset=["ts"]).sort_values("ts") +print(df.set_index("ts").resample("1D").size().describe()) +# the 38 all carry a trailing " (edited)" in the ts field — a scraper bug, recoverable +``` + +Thirty-eight rows lost to a scraper artefact, now visible in a file with their URLs, so they can be +re-collected rather than quietly vanishing. Had `format="mixed"` been omitted, the count would have +been in the thousands and would have looked equally clean. + +```python +# 6. the actual claim: who posts within a minute of whom +cols = ["ts", "actor", "id"] +pairs = pd.merge_asof( + df.loc[:, cols].rename(columns={"actor": "a", "id": "id_a"}), + df.loc[:, cols].rename(columns={"actor": "b", "id": "id_b"}), + on="ts", tolerance=pd.Timedelta("60s"), direction="nearest", allow_exact_matches=False, +) +print(pairs.query("a != b").groupby(["a", "b"]).size().sort_values(ascending=False).head(10)) +# acct_f acct_k 94 +# acct_k acct_f 91 +# acct_f acct_m 12 +``` + +`acct_f` and `acct_k` land within sixty seconds of each other 94 times. That is the finding, and it +is a sentence with a number in it. The graph comes next only to show *where* those two sit, not to +establish that they are linked. + +```python +# 7. the graph, with metrics computed before anything is drawn +import networkx as nx +edges = pairs.query("a != b").groupby(["a", "b"]).size().reset_index(name="weight") +G = nx.from_pandas_edgelist(edges, "a", "b", edge_attr="weight", create_using=nx.DiGraph) +btw = nx.betweenness_centrality(G, normalized=True) +print(sorted(btw.items(), key=lambda x: -x[1])[:5]) +nx.set_node_attributes(G, btw, "betweenness") +nx.write_gexf(G, "case.gexf") +# acct_f 0.31, acct_k 0.28, acct_m 0.04, ... +``` + +```bash +# 8. publish the queryable dataset and the figure +datasette serve -i case.db -m metadata.yml -o +``` + +Then `case.gexf` into Gephi, sized by the `betweenness` attribute already on the nodes, degree-1 +fringe filtered out, labels on for `acct_f` and `acct_k` only. And the daily volume series into +Datawrapper, with the two-day collection gap on 5–7 March drawn as a shaded band and the source +line reading "12,386 posts collected 2026-02-01 to 2026-03-31; 38 timestamps unparsed; four company +name variants merged". + +What this run establishes: a corrected counterparty ranking, a reproducible cleaning record +including one rejected merge, and a quantified timing relationship between two accounts. What it +does not establish: that `acct_f` and `acct_k` are operated by the same person, or coordinated at +all — posting within a minute is consistent with coordination, with both reacting to the same +trigger, and with a scheduler neither of them controls. The chart cannot distinguish those, and +saying so is the finding's credibility. ## Broader catalogues diff --git a/src/content/sheets/osint/geolocation.md b/src/content/sheets/osint/geolocation.md @@ -4,9 +4,9 @@ description: "Work out where a photo was taken, and when, from shadows, sun posi category: osint subcategory: "Geospatial" tags: [osint, geolocation, chronolocation, shadows, verification] -tools: [suncalc, shadowmap, geohints, qgis] +tools: [suncalc, shadowfinder, pysolar, shadowmap, shademap, geohints, qgis, gdal, exiftool] difficulty: advanced -updated: 2026-09-28 +updated: 2026-10-04 references: - name: "Bellingcat's Online Investigation Toolkit" url: "https://bellingcat.gitbook.io/toolkit" @@ -24,9 +24,10 @@ references: ## What this covers -Placing an image on the earth and in time without metadata. This is the discipline OSINT is best -known for and the most labour-intensive thing in it: a hard geolocation is hours of work, not -minutes. +Placing an image on the earth and in time without metadata. The place comes from matching what is +in the frame against the world; the time comes from the sun, which is the one thing in the picture +whose behaviour is exactly calculable. A hard geolocation is hours of work, and the chronolocation +that follows it is arithmetic you can show someone. ## Method @@ -36,50 +37,367 @@ minutes. 2. **Narrow the region.** Those details constrain the country or region long before they give you a point. Pole and bollard design alone often gets you to a handful of countries. 3. **Find a searchable anchor.** A business name, a phone number on a van, a street name, a bus - route number. One readable sign collapses the search. -4. **Match terrain.** Mountain ridgelines are effectively fingerprints and are visible from far - away. Compare against a terrain viewer. -5. **Confirm with imagery.** Satellite and street view, ideally from the era of the photo. Check - building footprints, not just the general scene. -6. **Chronolocate.** Shadow direction and length plus a known position gives you a time of day and - a range of dates. + route number. One readable sign collapses the search; spend your effort on making one legible + before you spend it on scanning imagery. +4. **Match terrain.** Ridgelines are fingerprints and are visible from tens of kilometres away. + This is the step that works when there is no signage at all, and the step that fails on flat + ground. +5. **Confirm with imagery.** Satellite and street view from the era of the photo. Match building + footprints and the relative geometry of fixed objects, not the general vibe of the scene. +6. **Measure the shadow, then compute.** Direction gives you the sun's azimuth, the length-to-height + ratio gives you its elevation. Together they give a time of day and, usually, two candidate date + windows per year rather than one. +7. **Say what would falsify it.** A geolocation you cannot break is one nobody checked. Name the + feature that would have to be in the wrong place for your answer to be wrong. -## Shadow and sun work +The judgement calls: decide early whether you are doing signage work or terrain work, because they +need different tools and mixing them wastes hours. Decide whether the date is given or derived — +if you take the date from the caption and then use the sun to "confirm" the caption, you have +proved nothing. And decide whether a precise coordinate should be published at all, before you +find one. -Given a location and a date, the sun's position is exactly calculable — so a shadow in a photo -constrains the time it was taken, and if you know the time it constrains the date. +## The shadow arithmetic -| Tool | What it does | -| --- | --- | -| [SunCalc](https://suncalc.org/) | Sun position, shadow direction and length for any place, date and time. The workhorse. | -| [ShadowMap](https://shadowmap.org/) | 3D buildings with cast shadows rendered at a chosen time. | -| [ShadeMap](https://shademap.app/) | Global shadow simulation including terrain and trees. | -| [Shadow Finder](https://github.com/bellingcat/ShadowFinder) | Inverts the problem: given shadow length and a timestamp, maps every point on earth where that shadow is possible. | +A vertical object of height `h` casting a shadow of length `s` on level ground fixes the sun's +elevation angle: + +```text +tan(elevation) = h / s + +h = 5.00 m, s = 8.20 m +elevation = atan(5.00 / 8.20) = atan(0.60976) = 31.37 degrees + +check: 5.00 / tan(31.37 deg) = 5.00 / 0.60976 = 8.20 m +``` + +The shadow points directly away from the sun, so the sun's compass bearing is the shadow's bearing +plus 180: + +```text +shadow bearing measured off the image : 062 deg +sun azimuth : (062 + 180) mod 360 = 242 deg +``` + +Three things break this. The ground must be level — a shadow running downhill is longer than the +arithmetic expects and will push your elevation too low. The object must be vertical and you must +be measuring its true height, not its foreshortened height in the frame. And `elevation` here is +geometric; at elevations below about 5 degrees atmospheric refraction lifts the apparent sun by +roughly half a degree, which matters at sunrise and sunset and nowhere else. + +Then the part people get wrong: an azimuth and elevation pair does **not** identify one date. The +sun retraces its declination on the way out of winter and on the way back in, so almost every +pair matches two windows in the year, roughly symmetric about the solstice. Expect two answers and +rule one out with something other than the sun — foliage, snow, an event in the background, the +clothes people are wearing. + +## Key tools + +### SunCalc + +Web only, free, no account. Give it a coordinate and a date and it draws the sun's track for that +day with the numbers underneath: altitude, azimuth, and a **shadow length** readout for an object +height you type in. That last field is why it beats the generic ephemeris sites — you can go +straight from the measurement you took off the image to the time that produces it, without +converting anything. + +```text +1. Enter the coordinate in "Set Lat/Lon", or drag the map pin onto the spot. +2. Set month and date in the dropdowns. Note the timezone line (TZ) it reports; everything + downstream is in that zone, and it is the usual source of a one-hour error. +3. Drag the time slider and watch two readouts: "Azimuth" against your measured sun bearing, + and "Shadow length [m]" against the shadow you measured, with "at an object level [m]" + set to the object's real height. +4. When both match at once, you have a time. When only the azimuth matches, you have two + times -- morning and afternoon mirror each other, and length is what separates them. +5. "Reverse Calculation" works the other way: fix the shadow and ask for the time. +6. Capture the page URL. The state lives in the fragment -- + #/<lat>,<lon>,<zoom>/<YYYY.MM.DD>/<HH:MM>/<object height> -- so the link reproduces the + exact reading rather than just the tool. +``` + +The azimuth it reports is a compass bearing measured clockwise from north, which is what you +measured off the image, so no conversion is needed here. Read-offs from a slider are worth about a +minute of precision at best; quote a window, not a timestamp. And the whole result is conditional +on the coordinate — a 50 m position error barely moves the azimuth, but a wrong *date* moves it by +degrees per week away from the solstices, which is exactly where the two-candidate-window problem +bites. + +### ShadeMap and Shadowmap + +Also web only. These solve the opposite problem: when the shadow in your image is cast by a +building or a ridge rather than by the subject, you cannot measure an object height, so SunCalc's +arithmetic has nothing to chew on. These render the shadows that real geometry casts at a chosen +moment and you match the *pattern* instead. + +```text +ShadeMap (shademap.app) + - loads OpenStreetMap building footprints plus a terrain model, so both urban and + mountain shadows are simulated + - date and time slider at the bottom; the shadow polygons redraw live + - the sun-exposure analytics (hourly / daily / annual) are a solar-study feature, + not an investigative one -- ignore them for this work + - free for interactive use on the site; the paid tiers are for embedding the engine + in your own map, not for using it here + +Shadowmap (shadowmap.org) + - 3D buildings with cast shadows, better for dense city blocks + - set the date and time, then orbit the camera to the photographer's approximate + position and compare the shadow edges against the frame + - sign-in gates some of the time controls + +What to capture either way: a screenshot at the matched time, the coordinate, the +date, the stated timezone, and the building or ridge you matched on. +``` + +Building shadows are only as good as the building heights in OpenStreetMap, which are frequently +absent, guessed, or recorded for a structure that has since been extended. Terrain shadows are more +trustworthy because the DEM is measured. Neither knows about trees, awnings, scaffolding or a +crane that was on site that month — so an unexplained shadow is a reason to doubt your position +before it is a reason to doubt the tool. + +### suncalc (Python) + +The same model as the website, as a library, which is what you want once you are checking more than +one candidate. Give it a coordinate and a timestamp and it returns the sun's position; loop it over +a year and you get every moment that matches your measurement, which is the honest way to find both +date windows instead of stopping at the first one. + +```bash +pipx install suncalc # or: pip install suncalc pandas +``` + +```python +import math, datetime +import pandas as pd +from suncalc import get_position, get_times + +lat, lon = 38.7075, -9.1364 # candidate coordinate +height, shadow = 5.00, 8.20 # metres, measured off the image +shadow_bearing = 62.0 # degrees, measured off the image + +# the two numbers the image gives you +elevation = math.degrees(math.atan(height / shadow)) # 31.37 +sun_bearing = (shadow_bearing + 180) % 360 # 242.0 + +def sun(dt): + # suncalc returns radians, and its azimuth is measured from SOUTH toward west -- + # add 180 degrees to get the compass bearing you measured off the image + p = get_position(pd.Timestamp(dt), lon, lat) + return (math.degrees(p["azimuth"]) + 180) % 360, math.degrees(p["altitude"]) + +# sweep the whole year at 5-minute resolution: this is what exposes the second window +start = datetime.datetime(2026, 1, 1, tzinfo=datetime.timezone.utc) +for step in range(0, 365 * 288): + dt = start + datetime.timedelta(minutes=5 * step) + az, alt = sun(dt) + if abs(az - sun_bearing) < 1.0 and abs(alt - elevation) < 0.5: + print(f"{dt:%Y-%m-%d %H:%M} UTC az {az:6.1f} elev {alt:5.1f}") + +# sanity-check the day itself: if your candidate time is after dusk, the coordinate is wrong +t = get_times(pd.Timestamp("2026-03-22 12:00:00"), lon, lat) +print("sunrise", t["sunrise"], "solar noon", t["solar_noon"], "sunset", t["sunset"]) +``` + +Two conventions to hold onto, because getting either wrong silently produces a plausible, wrong +answer. Azimuth comes back in radians measured from south, so the `+ 180` above is not optional. +Altitude is geometric and ignores refraction. Pass timezone-aware UTC datetimes and convert to +local time at the very end, after you have checked whether the location observes DST on that date +— a shadow matched in March against a September local time is the most common way this analysis +goes quietly wrong. The package is thin and has not changed since 2023; that is fine for a model +of celestial mechanics, and it is not a sign of anything. + +### pysolar + +A second implementation, which is the point of using it: run the same coordinate and timestamp +through both and you are checking your arithmetic rather than one library's. It is also the more +natural fit when you want irradiance or want to sweep without pandas in the way. + +```bash +pipx install pysolar # or: pip install pysolar +``` + +```python +import datetime +from pysolar.solar import get_altitude, get_azimuth + +lat, lon = 38.7075, -9.1364 +when = datetime.datetime(2026, 6, 15, 12, 30, tzinfo=datetime.timezone.utc) + +# pysolar's azimuth is already a compass bearing, clockwise from north -- no offset needed +print("altitude", get_altitude(lat, lon, when)) # 56.80 +print("azimuth ", get_azimuth(lat, lon, when)) # 216.45 + +# cross-check against suncalc for the same instant: 56.79 / 216.36 +# agreement to a tenth of a degree means the inputs are right; a whole-degree gap +# almost always means one of the two got a naive datetime and assumed local time +``` + +`get_altitude` applies a refraction correction by default, which is why it sits a hundredth of a +degree off suncalc's geometric value near the zenith and further off near the horizon — at 5 +degrees elevation the two will differ by about half a degree, and pysolar is the one closer to what +a camera actually saw. Passing a naive `datetime` is the failure mode: it will be interpreted as +UTC and you will get an answer that looks fine and is hours wrong. -Shadow Finder is the one worth knowing about, because it works when you have *no* candidate -location: +### ShadowFinder + +Bellingcat's inversion of the problem, for when you have no candidate coordinate at all. Feed it an +object height, a shadow length and a UTC timestamp and it maps every point on earth where a shadow +of that ratio is possible at that instant — a band, not a point, but a band you can intersect with +everything else you know. + +```bash +pipx install shadowfinder # or: pip install shadowfinder + +# object height, shadow length, date, time -- POSITIONAL, in that order +shadowfinder find 5 8.2 2026-03-22 16:00:00 --time_format=utc + +# when you already have the sun's elevation rather than a height/length pair +shadowfinder find_sun 31.37 2026-03-22 16:00:00 --time_format=utc + +# the subcommands and their arguments, before you guess +shadowfinder find --help +shadowfinder find_sun --help + +# writes a PNG map into the current directory, named for its inputs: +# shadow_finder_20260322-160000-Utc_5_8.2.png +ls shadow_finder_*.png + +# the timezone grid it uses for local-time input is generated once and cached +shadowfinder generate_timezone_grid +``` + +The older `--object-height / --shadow-length / --date-time` flag form that circulates in tutorials +no longer exists — current releases take positional arguments under a `find` subcommand and will +print a bare usage line if you use the old syntax. The output band is wide, and it is a band of +*possibility*: it tells you where such a shadow could fall, not where this one did. It is at its +most useful when the timestamp is independently known — from a livestream, a broadcast, a +timestamped post — and least useful when you derived the timestamp from the same photo. + +### GeoHints + +Web only, free, and the reference for step one. It catalogues the regional giveaways by country: +bollards, utility poles, traffic lights, road markings, licence plates, house numbers, sign shapes, +road-numbering schemes, post boxes, driving side, even the vehicles Google mounts its cameras on. +Built for GeoGuessr, which is precisely why it is organised the way an investigator needs — by +visible object rather than by country. + +```text +1. Pick the object in your frame that is most likely to be regulated nationally. + Ranked by how much they narrow things: licence plate format > road sign shape and + font > utility pole construction > bollard > kerb and road marking > house numbers. +2. Open that object's page and scan the per-country plates side by side. You are + looking for a disqualifier as much as a match -- "not this country" eliminates + faster than "maybe this one". +3. Cross two independent objects before you commit to a region. A pole design shared + by six countries and a sign font shared by four may intersect at one. +4. Check the Signs subsection separately: back-of-sign colour, chevron style and + bus-stop design are all national and all usually visible in a frame that has no + readable text at all. +5. Record which object you matched on and the page you matched it against. "Bollard + type matches Portugal" is checkable; "looks Iberian" is not. +``` + +It is a crowd-built reference, so coverage is uneven — rich for Europe and the GeoGuessr-popular +countries, thin for much of Africa and Central Asia, and it lags on recent sign redesigns and plate +reissues. Nothing on it is evidence by itself: a match narrows the search space, and the actual +geolocation still has to end at a specific place on a specific image. + +### QGIS and gdal_viewshed + +Terrain is the fallback when there is no text in the frame. A ridgeline's silhouette depends only +on the observer's position, so if you can trace the horizon in the photo you can test candidate +positions against a DEM until one produces that profile. QGIS has carried a native **Elevation +Profile** panel since 3.26, which is the fastest way to do it interactively. + +```text +QGIS elevation profile, for matching a horizon +1. Load a DEM (SRTM, Copernicus GLO-30, or a national LiDAR product where one exists). +2. Layer Properties > Elevation > tick "Represent elevation surface". Without this the + profile panel will not see the layer and the panel looks broken. +3. View > Elevation Profile to open the panel. +4. "Capture curve": draw a line from your candidate camera position out along the + bearing the photograph faces, past the furthest ridge. +5. Compare the plotted skyline against the photo's horizon. Use the vertical + exaggeration control to match the photo's apparent relief, then set it back to 1 + before you quote any number off it. +6. Wrong candidate positions fail obviously -- a peak in the wrong order, a saddle + that should be hidden. That is the point: this step eliminates fast. +``` + +For the reverse question — "could this spot have been seen from there at all" — `gdal_viewshed` +answers it as a raster, which is better than eyeballing when you have many candidates: + +```bash +# binary visibility from an observer 2 m above the surface, out to 15 km +# -ox/-oy are in the DEM's own CRS units, so reproject the DEM to a metric CRS first +gdal_viewshed -md 15000 -ox 499000 -oy 4283000 -oz 2 dem_utm.tif viewshed.tif + +# the minimum height a target would need to be visible, rather than a yes/no mask +gdal_viewshed -om DEM -md 15000 -ox 499000 -oy 4283000 dem_utm.tif minheight.tif + +# accumulate over a grid of observers: where could a photographer have stood +gdal_viewshed -om ACCUM -md 15000 -ox 499000 -oy 4283000 dem_utm.tif accum.tif + +# target height matters when the subject is a mast or a roofline, not the ground +gdal_viewshed -md 20000 -tz 30 -ox 499000 -oy 4283000 dem_utm.tif mast_visible.tif +``` + +A DEM is a bare-earth or surface model depending on which one you downloaded, and the difference +decides the answer: a bare-earth model will happily tell you that you can see through a forest and +a city. 30 m postings smooth real ridgelines, so a profile that is close but not exact is expected +and is not a match. And viewshed output is geometric — it knows nothing about haze, which is what +actually limits how far a camera sees on the day in question. Reprojection, cropping and +hillshading of these rasters are covered on +[Maps, Satellite & Street-Level Imagery](/sheets/osint/maps-and-satellite-imagery). + +### exiftool + +Before any of the above, check whether the question is already answered. Most images stripped by +social platforms have nothing, but files received directly from a source, pulled from a cloud +drive, or lifted off a messaging app's media folder frequently still carry GPS. ```bash -pip install shadowfinder -shadowfinder --object-height 2 --shadow-length 3.5 \ - --date-time "2024-06-15 14:30:00" +brew install exiftool # or: apt install libimage-exiftool-perl + +# GPS as a decimal pair you can paste into a map, and nothing else +exiftool -n -GPSLatitude -GPSLongitude -GPSAltitude photo.jpg + +# the direction the camera was pointing, which turns a point into a line of sight +exiftool -GPSImgDirection -GPSImgDirectionRef photo.jpg + +# every timestamp in the file, including the offset that tells you the local zone +exiftool -time:all -OffsetTime -OffsetTimeOriginal -a -G1 photo.jpg + +# a whole directory into one CSV, to find which file in a set still has coordinates +exiftool -csv -n -GPSLatitude -GPSLongitude -DateTimeOriginal ./images > gps.csv ``` -You need the object's height and its shadow's length in the same units, and a timestamp. It returns -a band of possible latitudes. +A GPS tag is a claim made by the device, not a fact: it can be a cached fix from the last place the +phone had signal, it can be the location the file was edited rather than shot, and it can be +written by hand. Treat it as a strong lead that still has to be confirmed against the frame. Full +metadata, container and manipulation analysis lives on +[Image & Video Forensics](/sheets/osint/image-video-forensics) rather than here. + +## Shadow and sun tools at a glance + +| Tool | What it does | +| --- | --- | +| [SunCalc](https://www.suncalc.org/) | Sun position, azimuth, altitude and shadow length for any place, date and time. The workhorse. | +| [ShadowMap](https://shadowmap.org/) | 3D buildings with cast shadows rendered at a chosen time. | +| [ShadeMap](https://shademap.app/) | Global shadow simulation including terrain as well as buildings. | +| [Shadow Finder](https://github.com/bellingcat/ShadowFinder) | Inverts the problem: given shadow length and a timestamp, maps every point on earth where that shadow is possible. | +| [suncalc-py](https://github.com/kylebarron/suncalc-py) | The SunCalc model as a Python library, for sweeping many candidates. | +| [pysolar](https://pysolar.readthedocs.io/) | Independent solar-position implementation, for cross-checking. | ## Visual reference [GeoHints](https://geohints.com/) catalogues the regional details that narrow a location — bollards, -traffic lights, road markings, utility poles, licence plates — by country. It was built for -GeoGuessr and is genuinely the best reference for this step. +traffic lights, road markings, utility poles, licence plates — by country. [Bellingcat's OpenStreetMap Search](https://osm-search.bellingcat.com/) and -[Spot](https://spot.bellingcat.com/) both let you search for *relationships* between features — +[Spot](https://www.findthatspot.io/) both let you search for *relationships* between features — "a church within 200m of a bridge over a river" — which is how you turn a described scene into -candidate coordinates. - -Further mapping, satellite and street-level tools are listed on +candidate coordinates. The Overpass queries underneath that idea are on [Maps, Satellite & Street-Level Imagery](/sheets/osint/maps-and-satellite-imagery). ## Tool reference @@ -96,13 +414,74 @@ Further mapping, satellite and street-level tools are listed on - **Satellite imagery has a date.** A building present in the photo and absent from imagery may simply be newer, or demolished. Check the capture date and look for historical imagery. -- **Terrain matching defeats you at low elevation.** Flat terrain has no ridgeline to match. -- **Shadow work needs the true timezone**, including DST, and the analysis is only as good as your - measurement of the shadow. +- **Terrain matching defeats you at low elevation.** Flat terrain has no ridgeline to match, and a + 30 m DEM will not reproduce a 10 m rise. +- **Shadow work needs the true timezone**, including whether DST was in force on that date in that + country, which is a different question from whether it is in force now. +- **One azimuth-and-elevation pair means two date windows.** Reporting the first one you find, and + not the second, is the most common way a chronolocation is wrong while being arithmetically + correct. +- **Circular reasoning.** If the date came from the caption, the sun cannot confirm the caption. + It can only confirm that the caption is internally consistent, which is a much weaker statement. - **Publishing a precise location endangers people.** For conflict imagery especially, consider whether the coordinates need to be public. - **Plausible is not confirmed.** A location that fits is a hypothesis. Confirmation means a - specific feature matching in a specific place. + specific feature matching in a specific place, named, so that someone else can go and disagree. + +## Worked example + +One datum: a photograph of a public square. No metadata, no caption beyond a claim that it was +taken "in September". A lamp standard in the foreground casts a clean shadow across flat paving. + +1. **Inventory and narrow.** The bollards, the kerb profile and the black-on-white house numbers + go into GeoHints object by object. The sign font and the bollard type intersect on Portugal; + a partial bus-stop livery in the background does not contradict it. +2. **Anchor.** A shopfront name is legible after a crop and upscale. Searching the name plus + "Lisboa" returns one address on a named square, and Street View from the same corner reproduces + the building line, the arcade spacing and the lamp standard. Candidate coordinate: + **38.7075, -9.1364**. +3. **Measure.** The lamp standard matches a documented 5.00 m type and its shadow, scaled against + the paving slabs, runs 8.20 m. The shadow's bearing, taken off the satellite image along the + same paving joint, is 062 degrees. +4. **Convert.** `atan(5.00 / 8.20)` = **31.37 degrees** solar elevation. Sun bearing = + `(62 + 180) mod 360` = **242 degrees**, so the sun is in the west-south-west and this is an + afternoon photograph. +5. **Sweep the year.** The `suncalc` loop above, at the candidate coordinate, with a 1-degree + azimuth and 0.5-degree elevation tolerance, returns seven matching instants in all of 2026, in + two clusters: + +```text +2026-03-21 16:00 UTC az 241.6 elev 31.0 +2026-03-22 16:00 UTC az 242.0 elev 31.2 +2026-03-23 16:00 UTC az 242.4 elev 31.5 +2026-03-24 16:00 UTC az 242.8 elev 31.7 + +2026-09-20 15:45 UTC az 242.1 elev 31.8 +2026-09-21 15:45 UTC az 241.8 elev 31.5 +2026-09-22 15:45 UTC az 241.6 elev 31.1 +``` + +6. **Convert to local time, carefully.** Portugal moved to UTC+1 on 29 March 2026 and back on 25 + October. The March window is therefore **16:00 local**, and the September window is **16:45 + local**. Same sun, different clock, entirely because of a date the arithmetic knows nothing + about. +7. **Break the tie with something that is not the sun.** The trees on the square are in full leaf + and the café terrace is laid out — both argue September over the third week of March. The + caption's "September" is now consistent with the image rather than assumed by it, because the + sun was computed independently and the foliage was the tiebreaker. +8. **Cross-check the model.** The same instant through `pysolar` returns an elevation within a + tenth of a degree of `suncalc`'s, which confirms the inputs rather than the conclusion. + +What you can assert: a square identified by a named shopfront and reproduced building geometry, and +an afternoon sun position consistent with **20–22 September 2026, around 16:45 local time**, with +21–24 March as a rejected alternative and the reason for rejecting it stated. + +What would falsify it: a shopfront that was at a different address on that date, paving that was +relaid between the Street View capture and the photograph (which would break the shadow +measurement, not the location), a lamp standard of a different height, or any evidence that the +foliage argument is wrong. The measurement that carries the most risk is the 8.20 m — a 10 per cent +error there moves the elevation by roughly 2.5 degrees and the window by several days, so quote it +with that tolerance or do not quote a day at all. ## Broader catalogues @@ -111,7 +490,7 @@ Further mapping, satellite and street-level tools are listed on ## More tools -Further tools for this area from the OSINT Newsletter Tools Library (), excluding those already listed above. +Further tools for this area from the OSINT Newsletter Tools Library ([Geolocation and Maps OSINT](https://tools.osintnewsletter.com/tool-categories/geolocation-and-maps-osint)), excluding those already listed above. | Tool | What it does | | --- | --- | diff --git a/src/content/sheets/osint/image-video-forensics.md b/src/content/sheets/osint/image-video-forensics.md @@ -196,8 +196,9 @@ ffmpeg -i video.mp4 -vf "select='gt(scene,0.3)',metadata=print:file=scenes.txt" `-c copy` matters for anything you will publish or hand to someone else: re-encoding an excerpt destroys the compression history that any later analysis would use, and makes your excerpt -unfalsifiable in the wrong direction. `-fps_mode vfr` replaced `-vsync vfr` in ffmpeg 5; both are -accepted on current builds but only the former is documented. +unfalsifiable in the wrong direction. `-fps_mode vfr` replaced `-vsync vfr` in ffmpeg 5, and `-vsync` +was removed outright in ffmpeg 9 — it now fails with `Unrecognized option 'vsync'`. Any older +cheatsheet or Stack Overflow answer using it will not run. Scene-change detection finds cuts, and cuts in material claimed to be a single continuous recording are worth explaining. It also fires on pans, flashes and camera shake, so read the list as diff --git a/src/content/sheets/osint/osint-foundations.md b/src/content/sheets/osint/osint-foundations.md @@ -4,9 +4,9 @@ description: "How to run an open-source investigation without burning yourself o category: osint subcategory: "Foundations" tags: [osint, methodology, opsec, verification] -tools: [hunchly, obsidian, logseq] +tools: [hunchly, obsidian, logseq, multipass, lima, docker, curl, wget, dig, shasum] difficulty: beginner -updated: 2026-09-28 +updated: 2026-10-04 references: - name: "Bellingcat's Online Investigation Toolkit" url: "https://bellingcat.gitbook.io/toolkit" @@ -30,10 +30,11 @@ references: ## What this covers -The habits that decide whether an investigation holds up: how you look at a target without -telling them, how you record what you found so it survives the page being deleted, and how you -avoid deciding the answer before you have it. Tooling is the easy part of OSINT. This is the part -that separates a finding from a guess. +The habits that decide whether an investigation holds up: how you look at a target without telling +them, how you record what you found so it survives the page being deleted, and how you avoid +deciding the answer before you have it. Most of this sheet is discipline rather than tooling, +because most of it has no command — and where a tool genuinely exists, the commands below are the +ones that keep you from leaking. ## The rule that matters most @@ -44,19 +45,459 @@ that only three people have seen tells those three people someone is looking. Decide before you start: is this target likely to notice, and does it matter if they do? -## Collection hygiene +## Method -1. **Separate identity.** A research browser profile, or better a separate VM, with its own - accounts. Never the account you use for anything else. Expect platforms to ban research - accounts eventually and do not build anything you cannot lose. -2. **Capture as you go, not afterwards.** Anything interesting gets archived the moment you see - it. Pages disappear, get edited, or go private within hours of someone noticing attention. -3. **Record the URL, the timestamp, and how you got there.** A screenshot with no source and no - date is worth nothing. The path you took to a finding is part of the finding. -4. **Keep raw separate from conclusions.** One place for what you collected, another for what you - think it means. Conflating them is how an assumption becomes a fact three notes later. -5. **Write down what you looked for and did not find.** Negative results stop you re-running the - same dead end next week, and they are what an honest report needs. +1. **Write the question down before you collect anything.** One sentence, answerable, with a + stated standard of proof. "Is this company linked to that one" is not a question; "does any + public filing show a shared officer, address or beneficial owner" is. +2. **Decide your exposure budget first.** Which identity touches the target, from what address, + and what happens if that address is logged. Changing this mid-investigation is how an account + gets burned. +3. **Capture as you go, never afterwards.** Anything interesting is archived the moment you see it. + Pages disappear, get edited or go private within hours of someone noticing attention. See + [Archiving & Evidence Preservation](/sheets/osint/archiving-and-evidence). +4. **Record the route, not just the result.** URL, UTC timestamp, the query that surfaced it, and + what you clicked to get there. A screenshot with no source and no date is worth nothing, and the + path to a finding is the part you forget first. +5. **Keep raw separate from conclusions, in different files.** One place for what you collected, + another for what you think it means. Conflating them is how an assumption becomes a fact three + notes later and a published claim three weeks later. +6. **Log what you looked for and did not find.** Negative results stop you re-running the same + dead end next week, and they are what an honest report needs in order to describe its own limits. +7. **Name the thing that would prove you wrong, then go look for it.** If you cannot state what + evidence would change your conclusion, you are not investigating, you are assembling. +8. **Stop at the question you wrote down.** The collection will always offer you more. More about a + bystander, more about a family member, more that is interesting and none of your business. + +## Key tools + +### Hunchly + +The one tool that removes the discipline problem, because it captures continuously rather than when +you remember to. It is a browser extension that records every page you visit while a case is open: +the full HTML and resources, the URL, a UTC timestamp and a hash of the capture, into a local case +file that is full-text searchable. That hash-at-capture-time is the evidentiary point — it is the +same argument as the manifest pattern in +[Archiving & Evidence Preservation](/sheets/osint/archiving-and-evidence), applied automatically to +everything you looked at including the pages that turned out not to matter. + +Commercial, subscription, now owned by Maltego Technologies, with a 30-day trial that needs no card. +There is no CLI and no API: this is a GUI tool and the workflow is the product. + +```text +1. Install the extension, open the desktop app, and create a case BEFORE you + start looking. Captures outside an open case are not recorded. +2. Set the case name to your case ID, not to the target's name. The case file + leaks the target's name to anyone who sees your screen or your backups. +3. Toggle capture ON. The extension icon state is the only indication; check it + after every browser restart, because a silent-off session is unrecoverable. +4. Add "selectors" — the names, handles, domains and phone numbers you are + tracking. Hunchly then highlights them on every page you visit and logs + which page each one appeared on. This is the feature people underuse: it + catches a name in a footer you would never have read. +5. Tag pages as you go. An untagged case file of 4,000 pages is a search index, + not a narrative. +6. Export: the case export carries the captured pages, the hashes, the timestamps + and the selector hit log. Export at milestones, not just at the end — the + case file is a single local database and a single local database can corrupt. +7. Storage choice matters. Local keeps everything on your machine; the hosted + option puts your captures and therefore your whole research pattern on + someone else's infrastructure. Pick deliberately and write down which. +``` + +What it does not do: a Hunchly capture is a record that *you* saw that content at that time, with no +third-party attestation. It is excellent provenance for your own process and weak evidence against +a determined challenge, which is why anything load-bearing still gets a public-archive capture and +an independent timestamp. It also captures everything, including pages you visited by accident and +pages containing other people's personal data — treat the case file as sensitive material in its own +right. + +### Obsidian or Logseq: the case vault + +Local, plain-text, file-per-entity notes with links between them. The reason this beats a document +is that an investigation is a graph of entities, not a narrative, and the thing you need six weeks +in is "every note that mentions this phone number". Obsidian suits entity-per-note work where the +links between people are the finding; Logseq's daily journal and outliner suit chronology-heavy +work. Both store Markdown on disk, so the vault is greppable, diffable and `git`-able without the +application. + +Structure the vault so that provenance cannot be separated from content: + +```text +case-2026-014/ + 00-question.md the question, the standard of proof, the stop condition + 10-entities/ + person-a-khan.md one file per entity, named by role not by guesswork + company-northgate.md + account-acct_f.md + 20-raw/ collected artefacts, never edited + 2026-10-04T0407Z-post123.warc.gz + 2026-10-04T0407Z-post123.info.json + MANIFEST.sha256 + 30-findings/ + f01-post-edited-after-upload.md + 40-negative/ + not-found.md every search that returned nothing, with the query + 90-log.md append-only: what you did, when, from which identity +``` + +Every entity note carries its own provenance header, so a claim can never be read without its +source: + +```markdown +--- +entity: company-northgate +type: company +aliases: ["Northgate Ltd", "Northgate Limited", "northgate ltd"] +confidence: medium +first_seen: 2026-10-03 +last_checked: 2026-10-04 +--- + +## Established + +- Registered number 09876543, UK. Source: Companies House API, retrieved + 2026-10-04T04:07Z. Raw: `20-raw/ch-09876543.json` (sha256 4b1c8e...). + +## Reported but unverified + +- Described as "the trading arm" in the filing at para 17. Single source, + uncorroborated, author has an interest. Do not repeat as fact. + +## Ruled out + +- Not the same as Northgate Mining Ltd (number 07654321). Different officers, + different address, name similarity only. Checked 2026-10-04. + +## Open + +- Beneficial owner behind the BVI parent. UK PSC filing names the parent only. +``` + +The `Ruled out` section is the one people skip and the one that saves the investigation. An +explicitly rejected hypothesis stays rejected; an implicitly rejected one resurfaces in a month as +a half-remembered lead and sometimes as a published error. + +Three cautions. Do not use wiki-style double-bracket links if the notes might ever be published or +converted — they do not resolve outside the application, and a dead link in a published finding +reads as sloppiness. Sync services put your entire research graph on third-party infrastructure, so +if the vault syncs, know where to. And a vault is not an archive: the notes reference the artefacts, +and the artefacts need their own hashes and timestamps. + +### The research account + +A separate identity that touches targets, which you expect to lose. Platforms ban research accounts +eventually — for scraping, for viewing too many profiles, for geographic inconsistency, for nothing +at all — so build nothing on it you cannot walk away from, and never let it share anything with an +account you care about. + +There is no tool for this. There is a checklist, and the order matters: + +```text +BEFORE the account exists + [ ] Decide the persona's purpose. A plausible-but-empty account is more + suspicious than an obviously new one with a stated interest. + [ ] Email first, from a provider that does not require a phone number, and + never from the provider your real accounts use. + [ ] Phone number, if the platform demands one. A number you control and can + lose. Never your own; a reused number links the research account to you + permanently and silently, because platforms match on it across services. + [ ] Password manager entry with the case ID in the title, so the account is + findable and disposable as a unit. + +WHEN the account exists + [ ] Separate browser profile at minimum, separate VM if the target is + competent. Never the same profile as anything personal — shared cookies, + shared localStorage and shared autofill all link them. + [ ] Consistent timezone, language and locale between the account's stated + location and the browser's. A profile claiming Lisbon from an en-GB + browser on UTC+0 is a detectable mismatch. + [ ] No contact import. Ever. One accidental contact sync hands the platform + your real address book and the platform hands the target "people you + may know". + [ ] Age the account before using it. A day-old account viewing 200 profiles + is a rate-limit and a ban; a two-month-old one is a user. + +ONGOING + [ ] One account per investigation where the targets could plausibly compare + notes. Cross-contamination between cases is how one burn becomes three. + [ ] Log every target the account touched, so you can assess the damage when + it is eventually identified. + [ ] Expect it to be lost. Export anything you need from it as you go. +``` + +The legal and ethical line is not a technical question and the checklist does not answer it. +Creating an account usually breaches a platform's terms of service; in some jurisdictions and some +employment contexts it does more than that, and a persona that actively deceives a person — rather +than merely observing public content — is a different act from a passive research account. Know +which one you are doing, get it authorised in writing if you are doing it for anyone but yourself, +and note that nothing here makes impersonating a real person or organisation acceptable. + +### Isolation: a disposable VM + +A compromise or a deanonymisation should cost you a throwaway machine rather than your real one and +the identity attached to it. Containers and virtual machines are not interchangeable here: +a container shares the host kernel and, with default settings, the host network identity, which +makes it fine for running CLI tooling and wrong for browsing a hostile target. Browse from a VM. + +```bash +# full VM, Ubuntu, disposable: multipass on macOS and Windows +multipass find # 26.04 is the current LTS alias +multipass launch lts --name research-014 --cpus 2 --memory 4G --disk 20G +multipass shell research-014 +multipass stop research-014 && multipass delete research-014 --purge # gone + +# or Lima, the same idea with a declarative YAML template. +# Note the locator form: `template:` — the older `template://` is deprecated as of Lima 2.0 +limactl create --name=research-014 template:ubuntu-lts +limactl start research-014 +limactl shell research-014 +limactl delete --force research-014 + +# snapshot before you touch the target, so you can roll back to clean. +# multipass only snapshots a STOPPED instance, so stop it first +multipass stop research-014 +multipass snapshot research-014 --name pre-target +multipass start research-014 +# ...after the session, roll back +multipass stop research-014 +multipass restore research-014.pre-target --destructive +``` + +```bash +# containers: correct for CLI tooling, with the network identity made explicit +docker run --rm -it \ + --dns 1.1.1.1 \ + --cap-drop ALL --security-opt no-new-privileges \ + -v "$PWD/out:/out" \ + python:3.13-slim bash + +# route a container's traffic through a proxy you control, and nothing else +docker run --rm -it \ + -e ALL_PROXY=socks5h://host.docker.internal:1080 \ + -e HTTPS_PROXY=socks5h://host.docker.internal:1080 \ + -v "$PWD/out:/out" python:3.13-slim bash + +# a container that cannot reach the network at all, for handling a hostile file +docker run --rm -it --network none -v "$PWD/sample:/sample:ro" python:3.13-slim bash + +# check the VM's egress before you use it, not after +multipass exec research-014 -- curl -s https://am.i.mullvad.net/json +``` + +For network-level rather than machine-level isolation, Whonix routes an entire VM's traffic through +Tor at the gateway, so a misconfigured application inside it cannot leak around the proxy, and +Tails gives you an amnesic live system that forgets everything on shutdown. Both are the right +answer when the consequence of being identified is serious; both are heavy enough that people skip +them for routine work, which is a defensible trade as long as it is a decision rather than a +default. + +What isolation does not buy you: a clean VM behind a VPN is still identified by the account you log +into, the browser fingerprint you present and the timing of your activity. Isolation limits the +blast radius of a mistake. It does not make you anonymous. + +### Touching a target without leaking + +When you need the bytes from a target's own server rather than from an archive, go via the command +line rather than a browser, because a browser sends dozens of headers you did not choose and runs +code the target wrote. Here the honest version matters: **neither `curl` nor `wget` has a +`--no-referer` flag, because neither sends a `Referer` header unless you ask it to.** Verified, a +default `curl` request arrives at the server carrying only `Host`, `User-Agent` and `Accept`. + +```bash +# exactly what a default curl sends — check this yourself rather than trusting it +curl -s https://postman-echo.com/headers | jq . +# {"headers":{"host":"postman-echo.com","user-agent":"curl/8.7.1","accept":"*/*", ...}} +``` + +```bash +# headers only, no body: cheapest possible look, and it reveals the stack +curl -sI 'https://target.example/page' + +# response headers AND body, with the body discarded — some servers lie on HEAD +curl -s -o /dev/null -D - 'https://target.example/page' + +# present a plausible browser UA. A default "curl/8.7.1" in a small site's log +# is a flag that says "someone is scripting against us" +curl -s -A 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0 Safari/537.36' \ + 'https://target.example/page' -o page.html + +# send no User-Agent at all, which is a different and sometimes better choice +curl -s -H 'User-Agent:' 'https://target.example/page' -o page.html + +# a Referer only when you deliberately want the log to show a plausible path +curl -s -A 'Mozilla/5.0 ...' -e 'https://www.google.com/' 'https://target.example/page' + +# show the redirect chain without following it into something you did not expect +curl -sIL -w '%{http_code} %{url_effective}\n' -o /dev/null 'https://target.example/short' + +# through a SOCKS proxy, with DNS resolved AT the proxy — socks5h, not socks5 +curl -s --proxy socks5h://127.0.0.1:9050 'https://target.example/page' -o page.html + +# wget equivalents, for a recursive pull you want to keep polite +wget -U 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0 Safari/537.36' \ + --wait=2 --random-wait --limit-rate=200k \ + --page-requisites --adjust-extension 'https://target.example/page' +``` + +`socks5h` versus `socks5` is the one that bites people: with `socks5`, `curl` resolves the hostname +locally and only the TCP connection goes through the proxy, so your resolver — and therefore your +ISP and often your employer — sees exactly which host you looked up. With `socks5h` the proxy does +the resolution. The same distinction applies to `ALL_PROXY` in the container examples above. + +Two things no flag fixes. The TLS handshake carries the hostname in SNI unless the server supports +Encrypted Client Hello, so a passive observer on your network learns which site you contacted +regardless of these options. And a target that fronts its site with a CDN sees your request through +that CDN's logging, which is a second party you did not choose. + +### Leak checks: IP, DNS and WebRTC + +Run these before you touch a target, after every network change, and after every VPN reconnect. +The failure you are looking for is not "am I behind a VPN" — it is "does something on this machine +resolve or connect outside the tunnel", which is invisible until you measure it. + +```bash +# the address a web server sees +curl -s https://am.i.mullvad.net/json | jq '{ip, country, city, mullvad_exit_ip, organization}' +curl -s https://ifconfig.co/json | jq '{ip, country, asn_org}' + +# the address your DNS RESOLVER egresses from — this is the leak that matters. +# If this is your ISP while the line above is a VPN exit, DNS is outside the tunnel. +dig +short TXT o-o.myaddr.l.google.com @ns1.google.com + +# resolver IP, your apparent client IP, and the EDNS Client Subnet your resolver +# is handing to authoritative servers. An "ecs" line means your /24 is being +# disclosed to every nameserver you query. +dig +short TXT whoami.ds.akahelp.net + +# which resolvers this machine is actually using, whatever the VPN client claims +scutil --dns | grep nameserver # macOS +resolvectl status | grep -A2 'DNS Serv' # systemd-resolved + +# does a hostname resolve the same inside and outside the isolated VM +dig +short target.example +multipass exec research-014 -- dig +short target.example + +# mullvad's own CLI, if that is your provider: state, and whether DNS is leaking +mullvad status +mullvad dns get +``` + +WebRTC and browser fingerprinting have no CLI equivalent, because the leak is the browser's and +only a browser reproduces it: + +```text +https://browserleaks.com/webrtc does the browser disclose your real local + and public IP via ICE candidates, around + the proxy. This is the classic VPN leak and + it is a browser setting, not a network one. +https://browserleaks.com/dns which resolvers the browser actually used +https://www.dnsleaktest.com/ extended test; run it, not the standard one +https://coveryourtracks.eff.org/ how distinctive your fingerprint is. Read + the "one in N browsers" number, not the + pass/fail badge. +https://browserleaks.com/geo whether the page can get a precise location + from the OS rather than from the IP +``` + +Read the fingerprint result carefully, because the intuition is backwards: hardening a browser with +unusual settings and a long extension list makes it *more* identifiable, not less. A stock browser +in a stock VM is often the better disguise than a heavily customised one, and the only reliable +counter to fingerprinting is to look like a large crowd. + +What these checks cannot tell you: whether the target correlated your visit with something else. +Timing, a reused screen resolution, the same unusual font set, a session that starts every weekday +at 09:10 UTC — none of that shows up in a leak test, and all of it is linkable across identities. + +### Chain of custody + +The record that lets you say, later, that the file you are producing is the file you received, and +that nothing happened to it in between that you have not written down. It costs a minute per +artefact and it is the first thing attacked when a finding matters. + +```bash +# the moment an artefact arrives, before you open it +IN=~/cases/2026-014/20-raw +mkdir -p "$IN" +cp /Volumes/USB/clip.mp4 "$IN/" # copy, never move +shasum -a 256 "$IN/clip.mp4" | tee -a "$IN/MANIFEST.sha256" +chmod 444 "$IN/clip.mp4" # read-only: work on copies + +# the custody note, as a sibling file, written now and never edited +cat > "$IN/clip.mp4.custody" <<'EOF' +artefact: clip.mp4 +sha256: 4b1c8e... # from MANIFEST.sha256 +received_at: 2026-10-04T04:07:56Z # date -u +%FT%TZ +received_from: source S-3 (see 10-entities/source-s3.md), in person, USB +provided_as: claimed original off a phone; no chain before this point +handled_by: J. Investigator +first_action: hashed, set read-only, copied to 50-work/ for analysis +onward: none +EOF + +# every derived file records what it came from, so no copy is ever orphaned +shasum -a 256 50-work/clip-frame-0137.png >> 50-work/DERIVED.sha256 +printf '%s\tderived from\t%s\n' 'clip-frame-0137.png' 'clip.mp4 (4b1c8e...)' \ + >> 50-work/DERIVED.index + +# append-only activity log, one line per action, in UTC +printf '%s\t%s\n' "$(date -u +%FT%TZ)" 'extracted frame 00:01:37.5 with ffmpeg 8.0' \ + >> ~/cases/2026-014/90-log.md + +# verify the whole raw tree before you hand anything over +shasum -a 256 -c "$IN/MANIFEST.sha256" | grep -v ': OK$' +``` + +Then timestamp the manifest, which is what converts your own record into a third party's assertion +about time — `ots stamp` and the RFC 3161 route are in +[Archiving & Evidence Preservation](/sheets/osint/archiving-and-evidence). + +Four rules that are not negotiable. UTC everywhere, because a local timestamp in a report with +international sources is ambiguous and ambiguity is attackable. Copy rather than move, so the +original stays where it was. Write the note at the time, because a custody note reconstructed from +memory is exactly as reliable as it sounds. And record the gap honestly: if you do not know where +the file was before your source handed it to you, the note says "no chain before this point" rather +than nothing, because an unstated gap reads as a concealed one. + +### The negative-results log + +The cheapest high-value habit on this page, and the one almost nobody keeps. A searchable record of +every query that returned nothing stops you repeating dead ends, tells you where your coverage +actually ends, and is the only honest basis for a sentence like "no public record of X exists". + +```text +# 40-negative/not-found.md — append-only, one block per attempt + +## 2026-10-04T04:12Z — Companies House officer search +query: "Khan" + date of birth 1979-03 +scope: all UK registered companies, active and dissolved +result: 0 matches with that DOB month +means: not registered as a UK officer under that spelling. Does NOT mean + not an officer: Companies House shows DOB month and year only, and + transliteration variants were not tested. +next: re-run via Bellingcat Name Variant Search for Cyrillic variants + +## 2026-10-04T04:31Z — Sherlock, handle "acct_f" +query: sherlock acct_f --print-found +result: 3 hits, all confirmed unrelated third parties +means: handle is reused; it is not a usable pivot for this person +next: pivot on the profile photo instead, not the handle + +## 2026-10-04T05:02Z — Wayback, target.example/about +query: CDX, matchType=prefix, from=20240101 +result: no captures +means: nothing about absence of the page. The site serves a robots + directive that Wayback honoured at crawl time. +next: check archive.today and Ghostarchive before concluding anything +``` + +The `means:` line is the whole point and it is the line that gets dropped. "Zero results" is a fact +about a tool's coverage, and turning it into a fact about the world is the single most common +overreach in open-source work. A people-search service returning nothing means that service has +nothing. A registry returning nothing means that registry, under that spelling, on that date. + +Write the `next:` line too. A negative result with a stated next step is a lead; a negative result +without one is just a gap you will rediscover. ## Verification @@ -94,6 +535,125 @@ Two sources that both trace back to the same original post are one source. use it. - **Forgetting the human cost.** Publishing that someone can be located has consequences for them. Minimise what you expose beyond what the finding requires. +- **A VPN indicator is not a leak test.** The client says "connected" while the system resolver + still egresses from your ISP, and a browser can hand out your real address over WebRTC around any + tunnel. Measure the resolver and the browser separately, after every reconnect. +- **`socks5` where you meant `socks5h`.** The proxy carries the connection but your own resolver + does the lookup, so the hostname is logged locally even though the traffic was not. +- **Hardening a browser makes it more identifiable.** An unusual configuration and a long extension + list are a fingerprint. A stock browser in a disposable VM hides in a larger crowd. +- **Zero results is a fact about a tool, not about the world.** Record what the gap means and what + it does not, in the same note, at the time. + +## Worked example + +One datum: a tip-off email naming a company, `Northgate Minerals Trading`, and asserting it is a +front. Nothing else. The question is not "is it a front" — that is a conclusion looking for support. +This is the first hour, before any collection, and what it buys you. + +```text +# 1. write the question down. 00-question.md +question: Does any public record link Northgate Minerals Trading to + [named entity] through a shared officer, address or owner? +standard: two independent primary sources per link, or it is "reported, unverified" +out of scope: the tipster's motive; family members of any officer +stop when: the question is answered either way, or three named registries + have been checked and returned nothing +exposure: registry APIs and archives only. NO requests to any Northgate- + controlled domain from an attributable address until step 6. +``` + +The exposure line is the decision that constrains everything after it. Having written it down, the +first four steps are forced: only sources that do not touch the target. + +```bash +# 2. establish the egress before anything leaves the machine +curl -s https://am.i.mullvad.net/json | jq '{ip, country, organization, mullvad_exit_ip}' +dig +short TXT o-o.myaddr.l.google.com @ns1.google.com +# web-visible IP: a VPN exit, country NL +# resolver egress: 203.0.113.x — the SAME ISP range as the host, not the VPN +``` + +DNS is outside the tunnel. Every hostname you look up is visible to the ISP, and will be whether or +not the HTTP request goes through the VPN. Fix that before step 3, not after: the resolver leak is +the one that persists, because it is a system setting and the VPN client reported "connected". + +```bash +# 3. isolation, and verify it from inside rather than trusting the launch +multipass launch lts --name research-014 --cpus 2 --memory 4G --disk 20G +multipass exec research-014 -- curl -s https://am.i.mullvad.net/json | jq -r .ip +multipass exec research-014 -- dig +short TXT o-o.myaddr.l.google.com @ns1.google.com +# both now report the VPN exit. Snapshot clean before use: +multipass stop research-014 && multipass snapshot research-014 --name pre-target +multipass start research-014 +``` + +```text +# 4. open the Hunchly case BEFORE the first search, not after +case name: 2026-014 (the case ID, never "Northgate") +selectors: Northgate Minerals Trading + Northgate Ltd + 09876543 + [officer surname, once known] +capture: ON — confirmed by the extension icon after the browser restart +``` + +Every page from here on is captured, hashed and timestamped whether or not you thought it mattered +at the time. That is the whole reason the case opens before the searching: the page you will need is +the one you skimmed on the way to something else. + +```bash +# 5. collect from sources that do not touch the target, and log the negatives +# (the registry and sanctions commands themselves live on the companies sheet) +printf '## %s — Companies House name search\nquery: "Northgate Minerals Trading"\nresult: 2 hits, gb/09876543 active + vg/1654321 (source dated 2019)\nmeans: a UK and a BVI company share the name. Not yet evidence they are related.\nnext: PSC filing on 09876543\n\n' \ + "$(date -u +%FT%TZ)" >> ~/cases/2026-014/40-negative/not-found.md +``` + +```bash +# 6. the only step that touches the target, and only after deciding to +# headers first: it answers the hosting question without fetching the page +multipass exec research-014 -- \ + curl -sI -A 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0 Safari/537.36' \ + 'https://northgate-minerals.example/' | head -12 +# Server: cloudflare → the origin IP is not in this response, and a CDN +# operator now has a log line for this request. Noted in 90-log.md. +``` + +A `HEAD` from a VM behind a VPN, with a browser UA, after an explicit decision — rather than a +browser tab opened reflexively in hour one from your own address. The difference costs ninety +seconds and is the difference between the target knowing and not knowing. + +```bash +# 7. the artefacts, hashed and custody-noted at the moment they arrive +IN=~/cases/2026-014/20-raw; mkdir -p "$IN" +multipass transfer research-014:/home/ubuntu/ch-09876543.json "$IN/" +shasum -a 256 "$IN/ch-09876543.json" | tee -a "$IN/MANIFEST.sha256" +chmod 444 "$IN/ch-09876543.json" +printf '%s\t%s\n' "$(date -u +%FT%TZ)" 'retrieved CH filing 09876543 via API from research-014' \ + >> ~/cases/2026-014/90-log.md +``` + +```text +# 8. the entity note, written as three separate claims. 10-entities/company-northgate.md +Established: number 09876543, UK, active. Source: CH API, 2026-10-04T05:12Z, + raw ch-09876543.json (sha256 4b1c8e...). +Reported/unverified: "a front". Single source, the tipster, who has an interest. + Not repeated as fact anywhere else in the case. +Ruled out: not Northgate Mining Ltd (07654321) — different officers, + different address, name similarity only. Checked 2026-10-04. +Open: beneficial owner behind the BVI parent. +``` + +After an hour the case holds: a written question with a stop condition, a fixed and verified +exposure posture, a continuous capture log, two registry records with hashes and custody notes, one +explicitly rejected lookalike company, and a negative-results entry that says what "two hits" does +and does not mean. + +What none of that establishes: whether Northgate is a front. The tipster's claim is still exactly +one uncorroborated assertion, recorded as such, in a section of the note that cannot be mistaken +for a finding. The substantive work now moves to +[Company & Financial Records](/sheets/osint/companies-and-finance) — but it moves there on top of a +record that will survive someone attacking it, which is the only thing this sheet is for. ## Broader catalogues diff --git a/src/content/sheets/osint/reverse-image-search.md b/src/content/sheets/osint/reverse-image-search.md @@ -4,9 +4,9 @@ description: "Find where an image came from and whether it predates the event it category: osint subcategory: "Images & Video" tags: [osint, images, verification, reverse-search] -tools: [google-lens, tineye, yandex, invid] +tools: [google-lens, tineye, yandex, bing, invid, ffmpeg, imagemagick] difficulty: beginner -updated: 2026-09-28 +updated: 2026-10-04 references: - name: "Bellingcat's Online Investigation Toolkit" url: "https://bellingcat.gitbook.io/toolkit" @@ -24,72 +24,361 @@ references: ## What this covers -Establishing whether an image is what it claims to be, by finding earlier copies. This is the +Establishing whether an image is what it claims to be, by finding earlier copies of it. This is the single highest-value check in visual verification and it takes under a minute, so it goes first, -always. Most viral misinformation is old footage relabelled. +always — most viral misinformation is old footage relabelled, and the relabelling is what a reverse +search catches. The engines disagree constantly, which is not a flaw to work around but the reason +you run more than one. ## Method -1. **Run several engines.** They index different corpora and disagree constantly. One engine - returning nothing means nothing. -2. **Sort by oldest**, not by relevance. You want the earliest appearance, which is the likely - original. -3. **Crop and re-run.** Engines match on the whole frame. Cropping to a distinctive element — a - sign, a building, a vehicle marking — often finds matches the full frame misses. -4. **For video, search keyframes.** Extract frames and search each; a video is only as findable as - its most distinctive still. -5. **Check the earliest hit's own context.** The oldest copy you can find is not necessarily the - original — read its caption and follow its sourcing. +1. **Run several engines on the unmodified file.** They index different corpora. One engine + returning nothing means that engine has nothing, and nothing more than that. +2. **Sort by oldest where the engine lets you.** Relevance ranking is actively unhelpful here — + you want the earliest appearance, not the most popular one. +3. **Preprocess and re-run.** Crop to the distinctive object, remove the overlay, flip it + horizontally, upscale it. Each of these is a different query against the same corpus and each + finds things the others miss. +4. **For video, extract keyframes and search each.** A video is only as findable as its most + distinctive still. +5. **Read the earliest hit's own context.** The oldest copy you can find is not necessarily the + original. Follow its caption and its sourcing before you call it the source. +6. **Archive what you found, with its date, before you cite it.** Search results are not stable — + see [Archiving & Evidence](/sheets/osint/archiving-and-evidence). -## Engines, and what each is good for +The judgement call that matters: decide what you are actually asking. "Where did this image come +from" and "does this image predate the event it is captioned with" are different questions, and +only the second one is usually answerable. The second is also the one that settles arguments, so +aim at it. -| Engine | Strength | -| --- | --- | -| [Google Lens](https://lens.google.com/) | Best at objects, text in images, landmarks and products. Strong OCR. | -| Yandex Images | Consistently the best at faces and at Eastern European and Central Asian content. Frequently finds what Google misses. | -| [TinEye](https://tineye.com/) | Oldest-first sorting and exact-match focus. The best tool for "when did this first appear". | -| Bing Visual Search | Good regional coverage; worth running as a third opinion. | -| [Search by Image](https://addons.mozilla.org/en-US/firefox/addon/search_by_image/) | Browser extension that fires one image at many engines at once. | -| [RootAbout](https://rootabout.com/) | Reverse image search across Internet Archive holdings. | +## Key tools + +### Google Lens + +Web only, free, no account needed for a one-off. The strongest engine for *objects* — products, +landmarks, vehicles, plants, architecture — and it has the best OCR of any of them, which means it +will often read a sign in your image and search the text without being asked. That is exactly what +you want for a photo with a shopfront in it, and exactly what you do not want for a photo whose +text is incidental. + +```text +1. lens.google.com, or the camera icon in Google Images. Upload the file; do not + paste a social-media URL, which searches the page rather than the image. +2. Drag the crop handles immediately. Lens defaults to the whole frame and the whole + frame is almost always the wrong query. Crop to one object and re-run. +3. Switch between the result modes. "Exact matches" is the one that answers the + verification question; the visual-similarity feed is for identifying the object. +4. "Find image source" (where offered) gives a dated list of pages carrying the image. + This is the closest Lens gets to TinEye's oldest-first view, and it is worth + scanning to the bottom rather than the top. +5. Capture: the result page as a screenshot plus the URLs of the earliest pages, and + the crop you actually searched -- a result is not reproducible without the crop. +``` + +You can also hand it a URL directly, which is useful for scripting a triage pass over a set of +images already hosted somewhere: + +```bash +# searches the image at that URL; the image must be publicly reachable +open "https://lens.google.com/uploadbyurl?url=https://example.org/photo.jpg" + +# url-encode anything with query parameters or the search will be of the wrong file +python3 -c 'import sys,urllib.parse; print("https://lens.google.com/uploadbyurl?url=" + urllib.parse.quote(sys.argv[1], safe=""))' \ + "https://example.org/img?id=123&size=full" +``` + +Lens personalises, so two people searching the same image get different result sets and neither is +reproducible from the other's screenshot. It is also weak on faces by design — it will refuse or +deflect — and weak on anything that is mostly sky, water or crowd. Dates shown next to results are +the page's claimed date, not the image's, and content farms backdate. + +### Yandex Images + +Web only, free. The one that finds what Google misses, consistently enough that skipping it is a +real gap in a verification. It is markedly better at faces, at crops of faces, and at Eastern +European, Russian and Central Asian content of every kind — if the image has Cyrillic in it or is +plausibly from that part of the world, Yandex is your first engine rather than your third. + +```text +1. yandex.com/images -> the camera icon -> upload, or paste an image URL. +2. The results page splits into "Similar images" and "Sites containing this image". + The second list is the evidential one; the first is a lookalike feed and will + cheerfully show you a different person with the same haircut. +3. Use the size filter in the left rail. Filtering to larger-than-your-copy is a fast + way to find an earlier, less-compressed version, which is usually closer to source. +4. Re-run on a tight crop of a face or a patch, separately. Yandex handles partial + matches far better than the others and a crop frequently returns hits the full + frame does not. +5. Read the result page titles even when the thumbnails look wrong -- its indexing of + Russian-language forums and VK is better than anything else available. +``` + +The same URL pattern works if you want to fire it from a script: + +```bash +open "https://yandex.com/images/search?rpt=imageview&url=https://example.org/photo.jpg" +``` + +The face matching is good enough to be an ethical question rather than a technical one: it will +identify private individuals from a crowd shot. Decide before you run it what you will do with a +hit, and whether the person in the frame is a subject of your investigation or a bystander. Beyond +that, Yandex's similarity feed is the most confident and the most wrong of any engine — it will +return a visually similar image with total assurance, so never treat a "similar" result as a match +without comparing fixed detail yourself. + +### TinEye + +Web only. Free for interactive use with no account; the published free API tier is **100 searches +per day and 300 per week** for non-commercial use, with paid tiers above that. The index is smaller +than Google's and it does not do visual similarity at all — it does exact and near-exact matching +of the same image, including crops, resizes and recolours. + +That narrowness is the whole point: **TinEye's "Oldest" sort is the single strongest piece of +evidence available for "this image predates the event it claims to show."** If TinEye returns a +copy indexed in 2019 and the caption says the photo is from last week's protest, the caption is +false, and no amount of argument about context changes that. + +```text +1. tineye.com -> upload the file, or paste an image URL. +2. Change the sort from "Best Match" to "Oldest". This is the only control on the page + that matters for verification. The options are Best Match, Most Changed, Biggest + Image, Newest and Oldest. +3. Read the first result's crawl date and open the page it sits on. TinEye's date is + when it first saw the image at that URL, not when the image was made -- it is a + latest-possible-creation date, which is the useful direction. +4. "Most Changed" is the second-most-useful sort: it surfaces the versions that have + been cropped, overlaid or edited, which shows you how the image has been used. +5. "Biggest Image" finds the least-compressed copy, which is what you want to hand to + the other engines and to any forensic step. +6. Capture: the oldest result's URL, TinEye's stated date for it, the total match + count, and an archive snapshot of that oldest page before it moves. +``` + +Absence from TinEye is close to meaningless — its crawl is much narrower than Google's and it does +not index most social platforms, so a genuinely viral image can show zero results. The date is a +crawl date and a page can have been republished at a new URL, so an "oldest" of last month does not +mean the image is from last month. And it is defeated by heavy re-editing in a way the +similarity-based engines are not: a mirrored, re-captioned, re-encoded repost may not match at all, +which is why the preprocessing step below exists. + +### Bing Visual Search + +Web only now. The Bing Search APIs — including Visual Search — were **retired on 11 August 2025** +and pre-retirement endpoints return HTTP 410, so any tutorial or script that calls a Bing visual +search API is dead; Microsoft's replacement is a grounding feature inside Azure AI Agents rather +than an image-match endpoint. The web interface still works and is still worth a minute, because +its regional coverage differs from both Google's and Yandex's. + +```text +1. bing.com/images -> the camera icon in the search box. (bing.com/visualsearch now + redirects to a Microsoft marketing page, not the tool.) +2. Upload, paste a URL, or drag a file in. +3. Use the on-image crop box, which Bing exposes more prominently than Google does -- + it is the best of the three for "search just this one object in the picture". +4. "Pages with this image" is the evidential list; "Related content" is not. +5. Capture the same way as the others: screenshot, URLs, and the crop you used. +``` + +Scriptable as a URL, with the usual caveat that it is an undocumented web parameter rather than a +supported interface and may change without notice: + +```bash +open "https://www.bing.com/images/search?view=detailv2&iss=sbi&q=imgurl:https://example.org/photo.jpg" +``` -Yandex's face matching is good enough to be an ethical question, not just a technical one. Think -about what happens to the person in the photo if you identify them. +No oldest-first sort, no date on most results, and a strong pull toward commercial and stock +imagery. Treat it as a third opinion that occasionally produces the one hit nobody else had, +rather than as a primary tool. -## Video keyframes +### Preprocessing with ImageMagick -[InVID / WeVerify](https://www.invid-project.eu/tools-and-services/invid-verification-plugin/) is -the standard browser plugin: it extracts keyframes from a video, runs them through several reverse -image engines, and also surfaces upload metadata and a magnifier for detail work. +The step that separates a reverse search that works from one that returns nothing. Engines match on +what is visually dominant, so an overlaid caption, a platform watermark or a wide establishing shot +all push the match toward the wrong thing. Each transformation below is a *separate query* — run +them all, do not pick one. -Manually, with ffmpeg: +```bash +brew install imagemagick # or: apt install imagemagick + +# what you are actually working with: dimensions tell you how much has been lost already +magick identify -verbose photo.jpg | head -40 + +# crop to the distinctive object. percentages, or pixels as WxH+X+Y +magick photo.jpg -crop 40%x40%+30%+20% +repage crop-sign.jpg +magick photo.jpg -crop 640x480+120+80 +repage crop-vehicle.jpg + +# mirror it. reposts are flipped constantly to defeat exact matching, and the engines +# do not try both orientations for you +magick photo.jpg -flop flipped.jpg + +# cut off a burned-in caption bar or platform watermark along the bottom +magick photo.jpg -gravity south -chop 0x90 nobar.jpg + +# upscale a small crop so an engine has pixels to work with. lanczos then a light +# unsharp is the combination that does not invent edges +magick crop-sign.jpg -filter Lanczos -resize 300% -unsharp 0x1 crop-up.jpg + +# pull detail out of a dark or flat region before cropping to it +magick photo.jpg -auto-level -sigmoidal-contrast 3,50% enhanced.jpg + +# strip metadata from anything you are about to upload to a third-party engine -- +# you are handing them the file, and the file may carry a source's GPS +magick photo.jpg -strip clean.jpg +``` + +`+repage` after a crop is not optional: without it the file keeps the original canvas geometry and +some tools will re-expand it. Upscaling adds no information — it makes a small crop palatable to an +engine's minimum-size requirements, and anything it appears to reveal is interpolation, so never +read a plate number off an upscale. And every enhancement you apply is a change to evidence: keep +the original untouched, work on copies, and record the exact command alongside the result, the same +way [Image & Video Forensics](/sheets/osint/image-video-forensics) handles hashing and the chain +from received file to analysed file. + +### ffmpeg + +Video does not reverse-search. Frames do. Pull the frames that are worth searching and you have +turned an unsearchable clip into eight or ten image queries. ```bash -# scene-change keyframes — the frames worth searching -ffmpeg -i video.mp4 -vf "select='gt(scene,0.3)'" -vsync vfr keyframe-%03d.jpg +brew install ffmpeg # or: apt install ffmpeg + +# every encoded keyframe, decoded fast because P- and B-frames are skipped outright. +# this is the cheapest first pass and usually enough +ffmpeg -skip_frame nokey -i video.mp4 -fps_mode vfr -q:v 2 key-%03d.jpg + +# frames at scene changes: 0.3 is a sane threshold, lower it to 0.1 for static footage +ffmpeg -i video.mp4 -vf "select='gt(scene,0.3)'" -fps_mode vfr scene-%03d.jpg + +# one frame every five seconds, as a fallback for a single unbroken shot that has no +# scene changes to detect +ffmpeg -i video.mp4 -vf fps=1/5 every5-%03d.jpg + +# a contact sheet, to pick the searchable frame by eye instead of opening fifty files +ffmpeg -i video.mp4 -vf "select='gt(scene,0.1)',scale=320:-1,tile=4x4" -fps_mode vfr sheet.png + +# the exact frame at a timestamp, once you know which second carries the readable sign +ffmpeg -ss 00:01:12.500 -i video.mp4 -frames:v 1 frame.png -# one frame every five seconds, as a fallback -ffmpeg -i video.mp4 -vf fps=1/5 frame-%03d.jpg +# crop, upscale and sharpen that frame in one pass, ready to hand to an engine +ffmpeg -i frame.png -vf "crop=400:300:820:460,scale=iw*4:ih*4:flags=lanczos,unsharp=5:5:1.0" search-me.png -# a contact sheet for eyeballing the whole video at once -ffmpeg -i video.mp4 -vf "select='gt(scene,0.2)',scale=320:-1,tile=4x4" -vsync vfr sheet.png +# mirror a frame, for the flipped-repost case +ffmpeg -i frame.png -vf hflip frame-flipped.png + +# blank out a burned-in platform logo rather than cropping the frame down +ffmpeg -i frame.png -vf "delogo=x=20:y=20:w=160:h=60" nologo.png ``` +`-vsync vfr` appears in most tutorials for this and has been **removed** in ffmpeg 9 — it now fails +with "Option not found". Use `-fps_mode vfr`, which replaced it in ffmpeg 5. Scene detection also +fires on camera pans and on cuts to black as readily as on a genuine change of location, so expect +a third of the frames to be useless; conversely a single-shot clip yields no scene frames at all, +which is why the `fps=1/5` fallback is there. Downloading the video in the first place, and reading +its container metadata, are on [Social Media Platforms](/sheets/osint/social-media-platforms) and +[Image & Video Forensics](/sheets/osint/image-video-forensics) respectively. + +### InVID / WeVerify + +A browser extension, free, from an EU research project. It collapses the whole video path above +into a few clicks and — the part that genuinely saves time — it extracts keyframes from a platform +URL without you downloading the file at all, then fires each keyframe at several engines in one go. + +```text +1. Install from weverify.eu/verification-plugin (Chrome and Firefox builds). +2. "Keyframes": paste the video URL. It fetches, segments and returns a grid of + keyframes with reverse-search buttons under each one. +3. Click through Google / Yandex / Bing / TinEye per keyframe. The buttons open the + engines in new tabs -- you still read the results yourself, it just saves the + upload step for each of the four. +4. "Magnifier": loads a still into a zoom-and-enhance panel, for reading a sign + without leaving the browser. +5. "Analysis": pulls the platform's own upload metadata for a YouTube/Facebook/X URL, + which gives you a latest-possible date for the upload -- not for the footage. +6. "Image forensics" runs ELA and noise filters. Use it for triage only; the real + version of that work is on the forensics sheet. +``` + +Platform integrations break whenever a platform changes its markup, so expect at least one tab of +the plugin to be dead at any given time; the keyframe extraction and the reverse-search buttons are +the durable parts. It gives you no reproducible command and no file hash, so anything you intend to +publish should be re-done with `ffmpeg` and recorded properly. And it uploads frames to third-party +engines, which is the same exposure question as any web tool — for sensitive material, extract +locally and decide deliberately what leaves your machine. + +### The smaller engines + +Worth a minute each when the big four come back empty, because they index corpora the others do not +touch. + +| Tool | What it is good for | +| --- | --- | +| [RootAbout](https://rootabout.com/) | Reverse image search across Internet Archive holdings — old web, scanned books and ephemera that no live-web crawler has. | +| [Search by Image](https://addons.mozilla.org/en-US/firefox/addon/search_by_image/) | Browser extension that fires one image at a configurable list of engines at once. The fastest way to run the full set without uploading four times. | +| [Karma Decay](http://karmadecay.com/) | Reverse search restricted to Reddit, which is where a surprising share of recycled images surfaces first. | +| [Pimeyes](https://pimeyes.com/) | Face search specifically, and paid beyond a teaser. Powerful, legally fraught in several jurisdictions, and a tool to think hard about before using on anyone who is not a subject. | + +None of these replaces the main four. Reaching for them is a sign that your query is wrong more +often than it is a sign that the image is unindexed — go back and crop harder before you go +further down this list. + ## Tool reference | Tool | What it does | Cost | | --- | --- | --- | | [InVID](https://weverify.eu/verification-plugin/) | A toolkit that supports the verification of videos and images. | free | +| [Google Lens](https://lens.google.com/) | Object, landmark and text matching with strong OCR. | free | +| [TinEye](https://tineye.com/) | Exact and near-exact matching with oldest-first sorting. | free / paid API | +| [RootAbout](https://rootabout.com/) | Reverse image search across Internet Archive holdings. | free | ## Pitfalls -- **Cropped, mirrored or filtered images defeat exact matching.** Flip the image horizontally and - re-run; recompressed and mirrored reposts are extremely common. -- **Absence of results is not originality.** It often just means the original is on a platform the - engine cannot index. -- **The oldest hit can still be a repost.** Read its context rather than treating the date as the - answer. -- **Screenshots of screenshots** lose the detail engines match on. Ask for the original file when - you can. +- **Cropped, mirrored or filtered images defeat exact matching.** Flip horizontally and re-run; + recompressed and mirrored reposts are extremely common and TinEye in particular will miss them. +- **Absence of results is not originality.** It usually means the original lives on a platform the + engine cannot index. Say "not found by X, Y and Z", never "original". +- **The oldest hit can still be a repost.** Read its context rather than treating the crawl date as + the answer. +- **Dates on result pages are the page's claim.** Content farms backdate, CMSs rewrite timestamps + on edit, and an archive snapshot is the only date you can stand behind. +- **Screenshots of screenshots** lose the detail engines match on. Ask for the original file. +- **Results are personalised and are not reproducible.** Archive the result page, and always record + the crop you searched — without it, nobody can repeat your query. +- **Uploading is disclosure.** Every web engine on this page keeps what you give it. For a file + from a source, strip it first and think about whether it should be uploaded at all. + +## Worked example + +One datum: a photograph circulating with the claim that it shows a named street during a protest +three days ago. + +1. **Unmodified file, four engines.** Lens returns stock-photo lookalikes. Bing returns news + aggregators from this week. Yandex returns a Russian-language forum thread. TinEye, sorted + **Oldest**, returns a first crawl dated **four years ago** on a regional news site. +2. **Open the oldest page.** It carries the same image, uncropped, with a caption naming a + different city and a different event. The claim is already broken at this point, and everything + after this is confirming rather than discovering. +3. **Confirm it is the same image, not a lookalike.** Crop both copies to the same corner — a shop + awning with a readable name — and compare. Identical awning, identical crack in the paving, + identical parked van. Fixed detail agreeing is the test; general resemblance is not. +4. **Explain why the other engines missed it.** The circulating version is mirrored and has a + caption bar burned along the bottom. `magick circulating.jpg -flop -gravity south -chop 0x90 + fixed.jpg` undoes both, and a re-run gets Lens to the same original — which demonstrates the + edit and shows the preprocessing step earning its place. +5. **Date the circulating version, not just the original.** TinEye's "Newest" sort and the + aggregator pages put the relabelled version's first appearance at four days ago, a day before + the protest it is captioned with — which is itself a finding. +6. **Archive before citing.** The four-year-old page, the aggregator pages, and the TinEye result + page all go to the Wayback Machine, because the regional news site is exactly the kind of + source that reorganises its URLs. See [Archiving & Evidence](/sheets/osint/archiving-and-evidence). + +What you can assert: the image was indexed on a named site four years before the event it is +captioned with, the circulating copy is a mirrored and cropped derivative of it, and specific fixed +details match between the two. + +What would falsify it: the two images being different photographs of the same unchanged street — +which is what the awning crack and the parked van rule out, and which is why you compare fixed +detail rather than overall appearance. A crawl date that TinEye got wrong would also do it, so the +archived copy of the four-year-old page, with its own publication date, is the thing worth having. ## Broader catalogues diff --git a/src/content/sheets/osint/social-media-monitoring.md b/src/content/sheets/osint/social-media-monitoring.md @@ -4,9 +4,9 @@ description: "Track accounts, hashtags and narratives across several platforms a category: osint subcategory: "Social Media" tags: [osint, monitoring, collection, social-media] -tools: [4cat, distill, snscrape] +tools: [4cat, zeeschuimer, changedetection.io, distill, yt-dlp, rsshub, rss-bridge, arctic-shift] difficulty: intermediate -updated: 2026-09-28 +updated: 2026-10-04 references: - name: "Bellingcat's Online Investigation Toolkit" url: "https://bellingcat.gitbook.io/toolkit" @@ -24,42 +24,444 @@ references: ## What this covers -Working several platforms at once, and watching things over time rather than looking once. Two -distinct jobs: **bulk collection** for later analysis, and **change detection** on pages you care -about. +Watching things over time instead of looking once, and holding what you collect so you can analyse +it later without re-collecting. Two distinct jobs with different tooling: **bulk collection** of a +corpus for analysis, and **change detection** on a small number of pages you care about. The third +thing, which is not optional, is doing either without the pattern of your collection becoming +visible to the people you are collecting on. -## Bulk collection +## Method -[4CAT](https://github.com/digitalmethodsinitiative/4cat) is the serious option — a self-hosted -capture-and-analysis platform that ingests from many platforms and ships analytical modules -(co-word networks, time series, image clustering) on top of what it collects. +1. **Decide what you are watching before you build anything.** A corpus for network analysis and an + alert on one bio edit need completely different infrastructure, and building the first when you + needed the second is the usual waste. +2. **Prefer the platform's own feed.** A native RSS or JSON endpoint is stable, keyless, polite and + does not attribute activity to an account. Reach for a scraper only when no feed exists. +3. **Store raw, analyse later.** Keep the untouched response alongside anything you derive from it. + You will want to re-run the analysis with a different question, and you will not want to + re-collect under a worse rate limit. +4. **Record gaps as gaps.** A failed poll is a hole in your data, not a quiet day. Log the failures + next to the results or your time series will lie to you. +5. **Set the cadence to the question.** Hourly is almost never justified. A daily poll that runs + for a year is more valuable, and far less conspicuous, than a five-minute poll that gets you + blocked in a week. +6. **Baseline before you conclude.** Coordination, surges and silences only mean something against + what normal looks like for that topic. + +The judgement calls: whether you need the content or only the fact that it changed — the second is +far cheaper and far quieter. Whether the collection needs an account at all, because the moment it +does, the collection has an identity and a history. And whether you are allowed to keep what you +are about to collect, which is a question to settle at the start rather than after you have a +database of it. + +## Collection hygiene + +Monitoring is repeated, scheduled, patterned contact with a target's content. One look is +invisible. The same request every fifteen minutes from one address for six months is a signature, +and on some platforms it is a signature attached to a logged-in account. + +```text +Account + - never your own, and never one that shares a recovery phone or email with your own + - a research account per investigation where the platform allows it; one burned + account should not cost you the others + - assume every logged-in query is retained and attributable, including search terms + - some platforms notify a user when a profile is viewed; know which before you look + +Rate + - daily is the default. Justify anything faster to yourself in writing + - randomise the interval. A poll at exactly :00 every hour is machine-obvious + - respect the stated limit, and treat a 429 as a signal to back off for hours, + not to retry in a loop + - set a real, honest User-Agent that identifies the project. It gets you unblocked + more often than a spoofed browser string does + +Network + - one address for all of it correlates every collection you run + - a residential or VPN exit is a trade: less rate limiting, more attribution risk + to whoever pays for it + - Tor is blocked by most of these platforms, so it is rarely the answer here + +Footprint + - log every request you make, with timestamp and response code. You need it to + distinguish "they stopped posting" from "we got blocked" + - decide retention up front. Bulk personal data attracts obligations even when + every item was public +``` + +The archiving and provenance side of this — hashes, snapshots, what makes a collected item citable +later — is on [Archiving & Evidence](/sheets/osint/archiving-and-evidence). Per-platform surfaces, +query syntax and the specific scrapers are on +[Social Media Platforms](/sheets/osint/social-media-platforms); this sheet is about running those +things repeatedly rather than once. + +## Key tools + +### 4CAT + +A self-hosted capture-and-analysis platform. It is the serious option because collection and +analysis live in the same place: the dataset you capture stays on your machine, and you can re-run +a different analysis over it months later without touching the platform again. That is the thing +no hosted service gives you. ```bash -git clone https://github.com/digitalmethodsinitiative/4cat -cd 4cat && docker compose up -d -# web UI is then served on port 80 +# the documented install is the compose file plus its .env, not a clone +mkdir 4cat && cd 4cat +curl -O https://raw.githubusercontent.com/digitalmethodsinitiative/4cat/master/docker-compose.yml +curl -O https://raw.githubusercontent.com/digitalmethodsinitiative/4cat/master/.env + +# the two settings worth reading before you start it: +# SERVER_BIND_ADDRESS=127.0.0.1 localhost only (the shipped default) +# PUBLIC_PORT=80 the port the web UI lands on +grep -E 'SERVER_BIND_ADDRESS|PUBLIC_PORT|DOCKER_TAG' .env + +docker compose up -d +docker compose logs -f backend # first start builds indexes; wait for it + +# the one-time admin link is printed in the logs, not emailed +docker compose logs frontend | grep -i 'create a new user\|token' +``` + +Creating a dataset, once it is up: + +```text +1. http://localhost:80 -> "Create dataset". +2. Pick the data source. Natively it collects 4chan, 8kun, Bluesky, Telegram, Tumblr + and TikTok (from a list of URLs). Everything else arrives as an upload. +3. Set the query and the date range. Date range is the field people leave open and + then wonder why the job runs for a day -- bound it. +4. Submit, then leave it. The backend queues and processes; the dataset appears in + your list with a status. Big Telegram or 4chan pulls take hours. +5. On the finished dataset, run processors rather than exporting immediately: + word frequencies, co-word networks, time series, top posters, image downloads + and clustering. Each produces a child dataset you can chain further processors on. +6. Export the one you want as CSV or NDJSON, and keep the parent dataset. The parent + is the thing you cannot re-create. +``` + +The platform coverage is the live constraint and it changes: X/Twitter, Instagram, LinkedIn, +Threads and Pinterest are no longer collected by 4CAT directly — they come in through Zeeschuimer +below, or as an upload from elsewhere. Datasets are big, and the Docker volumes will fill a disk +without warning, so check free space before a large pull. And a 4CAT capture is a snapshot of what +the platform served at that moment: edits and deletions after the capture are invisible to it, +which is a feature when you want the original and a trap when you assume it reflects the live site. + +### Zeeschuimer + +The companion extension, and the current answer to "how do I get X, Instagram or TikTok data into +4CAT". It watches the data the platform sends to your browser as you scroll, and keeps it. No API, +no scraping requests of its own — it records what your normal browsing already fetched. + +```text +1. Firefox only. Signed .xpi builds are on the project's releases page; Chrome is + not supported. +2. Install, then open the extension and enable the platform you are about to browse. + Supported: TikTok, Instagram, Threads, X/Twitter, Pinterest, Gab, Truth Social, + 9gag, Imgur, Douyin and RedNote/Xiaohongshu. +3. Browse normally -- a profile, a hashtag, a search. The item counter in the + extension ticks up as content loads. Content you never scrolled past is not + captured, because it was never sent to your browser. +4. Export as NDJSON or CSV, or put your 4CAT URL in the extension and upload straight + into it as a dataset. +5. Record how you browsed: which profile, which tab, how far you scrolled. That is + the sampling frame, and without it the dataset's coverage is undocumented. ``` -Self-hosting matters here: your dataset stays yours, and you can re-run analysis without -re-collecting. +This is manual collection with automatic recording, so it does not scale and it does not run +unattended — which is also why it survives platform changes better than scrapers do. It captures +through your logged-in session, so everything you collect is attributable to that account: use a +research profile in a separate Firefox container or profile, never your own. Platform support needs +constant maintenance and individual platforms do break; check the releases page before assuming the +extension is at fault rather than the site. + +### changedetection.io + +Self-hosted page watching. This is the one to run when the thing you need to notice is an edit — a +bio rewritten, a post deleted, an officer removed from a company page, a policy document quietly +amended, a price or a staff list changing. + +```bash +# docker, bound to localhost +docker run -d --restart always -p "127.0.0.1:5000:5000" \ + -v datastore-volume:/datastore \ + --name changedetection.io dgtlmoon/changedetection.io + +# or as a package, if you would rather not run a container +pip3 install changedetection.io +changedetection.io -d /path/to/empty/data/dir -p 5000 +``` + +```text +1. http://localhost:5000 -> "Add new watch", paste the URL. +2. Open "Edit" and set the filter before you set anything else. CSS selector, XPath, + JSONPath or jq -- point it at the narrowest element carrying the signal. Watching + a whole page means alerting on rotating ads, view counters and timestamps. + The "Visual Selector" tab picks the element by clicking it. +3. Set "Ignore text" for the lines you know churn: relative dates, "N views", + cookie-banner text. +4. Recheck time: per-watch. Daily for most things. The default applies to every watch + you add, so set it low once rather than per watch. +5. Notifications: Discord, email, Slack, Telegram or a webhook via Apprise. A webhook + into your own notes or ticketing is the one that leaves a record. +6. For a page that renders client-side, enable the Playwright/Sockpuppetbrowser + fetcher for that watch -- the default plain fetch sees an empty shell and will + report "no change" forever. +7. "Browser Steps" handles a page behind a form: click, fill, submit, then diff what + comes back. +8. Every change is stored as a snapshot with a diff view, which is the part that makes + it evidence rather than an alert. +``` + +Running it yourself means your watch list is not a third party's business record, which matters +when the watch list itself is sensitive. The cost is that a JS-heavy watch runs a real browser and +is far heavier than a text diff — a dozen of those on a small VPS will struggle. And a watch only +sees what an unauthenticated fetch from your server sees: a page that needs a login, or that +geo-varies, needs Browser Steps or will silently watch the wrong thing. + +### Distill + +The hosted equivalent, for when you will not run infrastructure. Same idea, less setup, and a free +tier that is genuinely usable for a handful of watches: **25 monitors total but only 5 in the +cloud, a 6-hour minimum cloud interval, 1,000 cloud checks a month, 2 devices, and 30 email alerts +a month**. Local monitors in the browser extension are unlimited but only run while the browser is +open. + +```text +1. Browser extension or distill.io. On the page you want, click the extension and + select the region -- it generates the selector for you. +2. Choose local (runs in your browser, unlimited, only while open) or cloud (runs + without you, capped as above). For anything that matters, cloud. +3. Set the check interval. Anything under 6 hours is a paid feature on the free tier, + and six-hourly is adequate for almost all of this work anyway. +4. Set the condition, not just "any change" -- Distill supports text conditions, so + "alert when the number changes" beats "alert when the page differs". +5. Alerts to email or webhook; the email allowance is the binding constraint on free. +6. Export the watch list as JSON periodically. It is the only part that is painful to + rebuild. +``` + +Your watch list and every snapshot live on their servers, which is the trade for not running +anything. For a target that could plausibly subpoena or compromise a third party, that is the wrong +trade and `changedetection.io` is the answer instead. The free tier's 6-hour floor also means you +will miss a post that goes up and comes down inside a window, which is precisely the kind of +deletion worth catching — if that is the scenario, self-host and poll faster. + +### yt-dlp with a download archive + +For recurring capture of a channel's output, the archive file is the whole trick: it records the ID +of everything already fetched, so the next run picks up only what is new. That turns a one-off +download into a monitor you can cron. + +```bash +pipx install yt-dlp + +# first run: establish the archive. --break-on-existing stops as soon as it meets +# something already recorded, so later runs walk only the new items +yt-dlp --download-archive archive.txt --break-on-existing --lazy-playlist \ + -o '%(upload_date)s-%(id)s.%(ext)s' \ + 'https://youtube.com/@channel/videos' + +# metadata-only monitoring: no video files, just the record that it existed +yt-dlp --download-archive seen.txt --break-on-existing \ + --skip-download --write-info-json --write-thumbnail \ + 'https://youtube.com/@channel/videos' + +# bound it by date instead, for a backfill of a known window +yt-dlp --dateafter 20260101 --datebefore 20260401 --download-archive archive.txt URL + +# subtitles, which turn a channel's output into greppable text as it arrives +yt-dlp --download-archive seen.txt --break-on-existing --skip-download \ + --write-auto-subs --sub-langs en 'https://youtube.com/@channel/videos' + +# throttle it. these two flags are the difference between a monitor and a nuisance +yt-dlp --sleep-requests 2 --sleep-interval 10 --max-sleep-interval 30 \ + --download-archive archive.txt URL + +# several channels in one run, each stopping at its own first-seen item +yt-dlp --break-per-input --break-on-existing --download-archive archive.txt \ + -a channels.txt + +# a cap, so a misconfigured run cannot pull a thousand files overnight +yt-dlp --max-downloads 50 --download-archive archive.txt URL +``` -## Change detection +The archive file is state: back it up, and never delete it to "start fresh" unless you mean to +re-download everything. `--break-on-existing` assumes the listing is newest-first, which it is for +channels and playlists and is not for some search result pages — on those, drop it and let the +archive do the skipping. Extractors break when platforms change their pages, so `pipx upgrade +yt-dlp` before blaming a URL, and a cron job that has silently failed for three weeks is worse than +no monitoring at all, so alert on non-zero exits. Adding `--cookies-from-browser` attaches a real +session to every request and makes the whole monitor attributable; avoid it unless the content +genuinely requires a login. Full metadata usage is on +[Social Media Platforms](/sheets/osint/social-media-platforms). -[Distill](https://distill.io/) watches a page or a page region and alerts on change. Good for -noticing when a target edits a bio, deletes a post, changes a company officer list, or quietly -updates a policy document. +### Native feeds, before you reach for a bridge -Point it at the narrowest element that carries the signal, not the whole page, or you will drown in -alerts from rotating ads and timestamps. +Several platforms still publish perfectly good feeds that nobody uses because everyone assumes RSS +died. These are keyless, stable, cheap to poll and attach to no account — the best monitoring +surface available, where it exists. + +```bash +# YouTube channel, by channel ID (the UC... form; @handles do not work here) +curl -s 'https://www.youtube.com/feeds/videos.xml?channel_id=UCXuqSBlHAE6Xw-yeJA0Tunw' + +# a subreddit's new posts +curl -s -A 'research-monitor/1.0 (contact: you@example.org)' \ + 'https://www.reddit.com/r/osint/new/.rss' + +# one Reddit user's activity +curl -s -A 'research-monitor/1.0 (contact: you@example.org)' \ + 'https://www.reddit.com/user/someuser/.rss' + +# any Mastodon account, on any instance +curl -s 'https://mastodon.social/@Gargron.rss' + +# a Bluesky profile +curl -s 'https://bsky.app/profile/bsky.app/rss' + +# a GitHub user's public activity, which dates account behaviour precisely +curl -s 'https://github.com/bellingcat.atom' + +# pull just the timestamps, to see cadence without reading content +curl -s 'https://www.youtube.com/feeds/videos.xml?channel_id=UCXuqSBlHAE6Xw-yeJA0Tunw' \ + | grep -oE '<published>[^<]+' | sed 's/<published>//' +``` + +Reddit rate-limits these hard and will return HTTP 429 to an anonymous, default-User-Agent client +within a handful of requests; a descriptive User-Agent and a gap of several seconds between calls +fixes it, and Reddit's `search.rss` endpoint is throttled more aggressively than the subreddit and +user feeds. YouTube's feed carries only the most recent entries, so it monitors but does not +backfill. Bluesky's profile RSS covers posts and not replies or likes — the AT Protocol endpoints +on [Social Media Platforms](/sheets/osint/social-media-platforms) go deeper. And X/Twitter publishes +nothing of this kind: `snscrape` has been non-functional against it since the 2023 access changes +and the public Nitter instances are largely dead, so there is no quiet monitoring surface for X at +all — only a logged-in session, with everything that implies. + +### RSSHub and RSS-Bridge + +For the platforms that killed their feeds. Both are self-hosted services that scrape a site and +re-emit it as RSS or Atom, so your monitor still speaks one protocol no matter how many platforms +are behind it. + +```bash +# RSSHub -- the larger of the two, ~1000 routes +docker run -d --name rsshub -p 1200:1200 diygod/rsshub +# or with the compose file, which adds redis caching and a browser for JS sites +wget https://raw.githubusercontent.com/DIYgod/RSSHub/master/docker-compose.yml +docker compose up -d + +# routes are paths. a Telegram public channel: +curl -s 'http://localhost:1200/telegram/channel/awesomeRSSHub' +# the route catalogue lives at docs.rsshub.app/routes/ + +# RSS-Bridge -- fewer routes (~450) but simpler, and a web UI that builds the URL +docker create --name=rss-bridge --publish 3000:80 \ + --volume $(pwd)/config:/config rssbridge/rss-bridge +docker start rss-bridge + +# bridges are query parameters rather than paths +curl -s 'http://localhost:3000/?action=display&bridge=<BridgeName>&format=Atom' +``` + +Self-host both. The public `rsshub.app` demo instance returns HTTP 403 to ordinary requests and is +not a reliable backend for anything you depend on, and a shared public instance makes your watch +list someone else's log file either way. Expect individual routes to break: they are scrapers +wearing an RSS hat, and when a platform changes its markup the route returns an empty feed rather +than an error — which reads exactly like "the target stopped posting". Check periodically that a +route still returns items, and never conclude silence from an empty bridge feed without confirming +against the site. + +### curl and jq on a schedule + +For a public JSON endpoint, a few lines beat any framework. The pattern that matters is the +watermark: store the timestamp of the newest item you have seen, and ask only for things after it. + +```bash +# poll Arctic Shift for new Reddit posts in a subreddit since the last run. +# full Arctic Shift usage is on the social media platforms sheet; this is the +# monitoring shape of it +AS='https://arctic-shift.photon-reddit.com/api' +STATE=~/monitor/last_seen_osint +SINCE=$(cat "$STATE" 2>/dev/null || echo 0) + +curl -sS --fail -G "$AS/posts/search" \ + --data-urlencode 'subreddit=osint' \ + --data-urlencode "after=$SINCE" \ + --data-urlencode 'sort=asc' \ + --data-urlencode 'limit=100' \ + -o /tmp/new.json || { echo "poll failed $(date -u +%FT%TZ)" >> ~/monitor/errors.log; exit 1; } + +# advance the watermark only on a successful fetch, or a failure silently +# becomes a gap you never notice +jq -r '.data[-1].created_utc // empty' /tmp/new.json | grep . && \ + jq -r '.data[-1].created_utc' /tmp/new.json > "$STATE" + +# append raw, then derive. the raw file is the thing you cannot regenerate +cat /tmp/new.json >> ~/monitor/osint-raw.ndjson +jq -r '.data[] | [.created_utc, .author, .title] | @tsv' /tmp/new.json +``` + +```bash +# a Bluesky author feed, keyless, as a second example of the same shape +curl -sS --fail 'https://public.api.bsky.app/xrpc/app.bsky.feed.getAuthorFeed?actor=bsky.app&limit=50' \ + | jq -r '.feed[] | [.post.indexedAt, .post.record.text] | @tsv' + +# crontab: daily, at a minute that is not :00, with failures mailed to you +# 37 6 * * * /home/you/monitor/poll-osint.sh >> /home/you/monitor/run.log 2>&1 +``` + +`--fail` is the flag that turns a silent HTML error page into a non-zero exit, and without it your +monitor will happily append a Cloudflare block page to its dataset for a month. Advance the +watermark only after a confirmed good response, log every failure with a timestamp, and treat the +log as part of the dataset — the difference between "they went quiet" and "we were blocked" lives +there and nowhere else. Note that Bluesky's `getAuthorFeed` is open without a token but +`searchPosts` on the same public host returns 403 unauthenticated, which is the sort of asymmetry +worth checking per endpoint before you build a schedule around it. + +### Arctic Shift for Reddit history + +The live successor to Pushshift, and the only practical way to get Reddit content that has since +been deleted or edited. For monitoring it plays a specific role: it is the backfill that gives your +forward-looking poll a baseline, and the thing you check when a post you captured disappears. + +```bash +AS='https://arctic-shift.photon-reddit.com/api' + +# the baseline: what this account did before you started watching it +curl -s -G "$AS/comments/search" --data-urlencode 'author=some_user' \ + --data-urlencode 'sort=asc' --data-urlencode 'limit=100' \ + | jq -r '.data[] | [.created_utc, .subreddit] | @tsv' + +# posting cadence by hour of day, which is what exposes a coordinated account +curl -s -G "$AS/posts/search" --data-urlencode 'author=some_user' \ + --data-urlencode 'limit=500' --data-urlencode 'fields=created_utc' \ + | jq -r '.data[].created_utc' \ + | python3 -c 'import sys,datetime,collections; c=collections.Counter(datetime.datetime.fromtimestamp(int(l), datetime.UTC).hour for l in sys.stdin); [print(f"{h:02d} {c[h]}") for h in range(24)]' + +# did a post you captured actually get removed, or did you lose it +curl -s -G "$AS/posts/ids" --data-urlencode 'ids=t3_abc123' | jq '.data[0] | {title, selftext, removed_by_category}' +``` + +Keyword search needs an accompanying `author`, `subreddit`, `link_id` or `parent_id` — bare keyword +sweeps are refused, and so are very broad author or subreddit queries. The archive holds what it +ingested at the time, so an edit made after ingestion is invisible and a post deleted within +seconds may never have been captured at all; a record here proves the text was published, not that +it is live. Bulk dumps for offline work are published separately from the API. Query syntax in full +is on [Social Media Platforms](/sheets/osint/social-media-platforms). ## Narrative and coordination analysis - **Posting-time clustering** exposes networks. Accounts that post within seconds of each other, - repeatedly, are coordinated. + repeatedly, are coordinated. The `jq` plus histogram pattern above is enough to see it. - **Identical phrasing across accounts** is the strongest single indicator of copy-paste campaigns. - **Follower overlap** between accounts is more telling than follower count. -- [The Information Laundromat](https://information-laundromat.com/) compares content across sites to - find where the same text is being syndicated. +- [The Information Laundromat](https://informationlaundromat.com/) compares content and technical + metadata — ad IDs, analytics tags, registration details, code structure — across sites, to find + where the same text is syndicated and which sites share infrastructure. Built by the Alliance for + Securing Democracy, which merged into the Institute for Strategic Dialogue on 1 January 2026; the + tool remains live at the address above. The older hyphenated domain no longer resolves. +- Graphing and clustering what you have collected is on + [Data Analysis & Visualisation](/sheets/osint/data-analysis-and-visualisation). ## Tool reference @@ -72,17 +474,73 @@ alerts from rotating ads and timestamps. | [Maltego Graph](https://www.maltego.com/downloads/) | Maltego Graph is an investigation platform that combines two things at once: (1) It acts as a search tool, and (2) It creates a graph establishing links… | partly free | | [Pinpoint](https://journaliststudio.google.com/pinpoint/about) | A tool by Google to catalogue uploaded documents and files, providing automated text recogntion, indexing, audiotranscriptions and other (AI-powered)… | free | | [Time.Graphics](https://time.graphics) | A tool for creating, visualizing, and managing timelines online. | partly free | +| [Zeeschuimer](https://github.com/digitalmethodsinitiative/zeeschuimer) | Firefox extension that records social media data as you browse it, for import into 4CAT. | free | +| [changedetection.io](https://changedetection.io/) | Self-hosted page-change monitoring with CSS/XPath/jq filters and diff history. | free | +| [Distill](https://distill.io/) | Hosted page-change monitoring, browser extension plus cloud checks. | partly free | ## Pitfalls +- **Dead tools still appear in tutorials.** `snscrape` and the public Nitter instances do not work + against X. Bing's Visual Search API and the rest of the Bing Search family were decommissioned in + August 2025. Check a tool is alive before you design a collection around it. - **Rate limits end collections mid-run.** Check for gaps before you analyse; a missing day looks like silence rather than a failure. +- **An empty feed is ambiguous.** A broken RSSHub route, a changed selector and a target who + stopped posting all look identical downstream. Monitor your monitors. - **Sampling bias reads as a finding.** If a tool only reaches accounts above some follower count, - its "network" is an artefact of that cutoff. + or only what you happened to scroll past, its "network" is an artefact of that cutoff. - **Coordination needs a baseline.** Fans of the same thing post about it at the same time. Compare against normal behaviour for the topic before calling it inauthentic. +- **Monitoring is contact.** Scheduled requests are a pattern; a logged-in monitor is an + attributable pattern. Decide what that costs before you start, not after. - **Storage and legality.** Bulk personal data attracts data-protection obligations even when every - individual item was public. + individual item was public, and a dataset is harder to delete than to collect. + +## Worked example + +One datum: a single Telegram channel name, from a screenshot, alleged to be seeding a story that +later appeared on a cluster of news-like websites. + +1. **Baseline before watching.** Pull the channel's existing history into 4CAT as a Telegram + dataset with a bounded date range. You now know its normal posting cadence and vocabulary, which + is what any later claim of a surge has to be measured against. +2. **Find the forward-looking surface.** Telegram has no public feed, so an RSSHub route + (`/telegram/channel/<name>` on your own instance) gives you a pollable endpoint. Verify it + returns items today, so that an empty feed next month means something. +3. **Schedule it honestly.** A daily `curl` with a stored watermark, a descriptive User-Agent, and + a failure log. Not hourly — a story that takes days to syndicate does not need fifteen-minute + resolution, and the slower poll survives longer. +4. **Watch the downstream sites for edits, not just posts.** Each suspected site gets a + `changedetection.io` watch filtered to the article body, so a quietly amended paragraph or a + removed byline raises an alert with a stored diff. This is the part that a collection-only + approach misses entirely. +5. **Capture the video claims properly.** The channel posts clips; `yt-dlp --download-archive` + with `--write-info-json` on its linked channels picks up only what is new each day and records + `upload_date` for each, giving every clip a latest-possible date. +6. **Test the syndication claim.** Feed one article URL to the Information Laundromat. Content + similarity tells you the text is shared; the technical indicators — a common analytics ID across + four of the sites — tell you something stronger, because wording can be copied by anyone and a + shared tracking ID usually cannot. +7. **Cross-check the Reddit leg.** Arctic Shift for posts linking those domains, grouped by + subreddit and by hour, shows whether the amplification is a handful of accounts on a schedule or + genuine spread. +8. **Archive as you go.** Every page you will cite goes to a snapshot at the time you saw it, per + [Archiving & Evidence](/sheets/osint/archiving-and-evidence) — these sites edit and disappear, + which is the behaviour you are documenting. + +What you can assert: a dated posting history for the channel, dated first appearances on each +downstream site, stored diffs of any subsequent edits, and a shared technical indicator linking +some of those sites. + +What would falsify it: an RSSHub route that broke silently mid-period, which would turn a real gap +into an apparent one — so the run log, showing a successful fetch every day, is doing as much +evidential work as the data. A shared analytics ID that turns out to belong to a common CMS +template, or to an agency that serves unrelated clients, would break the infrastructure link; check +what else carries that ID before leaning on it. + +## Broader catalogues + +- [Social Media OSINT](https://tools.osintnewsletter.com/tool-categories/social-media-osint) ## Sources diff --git a/src/pages/credits.astro b/src/pages/credits.astro @@ -134,9 +134,8 @@ SOFTWARE.`; <p> <a href="https://ired.team" target="_blank" rel="noopener">ired.team</a> — the red teaming notes of <strong>Mantvydas Baranauskas</strong> (<a href={iredSource.authorUrl} target="_blank" rel="noopener">@mantvydasb</a>) — - is the reference behind a number of sheets in Exploitation, Password Attacks, Privilege Escalation, - Tunneling &amp; Pivoting and DFIR. His write-ups on process injection, defense evasion and persistence are - some of the most careful hands-on documentation of those techniques anywhere. + is the reference behind sheets in Exploitation and DFIR. His write-ups on process injection and on + Windows internals are some of the most careful hands-on documentation of those subjects anywhere. </p> <p> <strong>ired.team publishes no licence.</strong> That grants no right to copy it, so nothing from it is