commit d1f3335bc7a7af813e54950b0f8b0027596f0317
parent 578fa3b5dc8ce50fd6216a3cc255e55e6b2ac9c0
Author: DAEMON <zer0sec.xp@icloud.com>
Date: Sat, 3 Oct 2026 18:51:29 +0100
Finish the OSINT usage pass: the last six sheets
Completes the OSINT tool-usage work. archiving-and-evidence,
data-analysis-and-visualisation, osint-foundations, geolocation,
reverse-image-search and social-media-monitoring were the remaining
catalogues — between 0 and 3 code blocks each, no per-tool sections.
They now follow usernames-and-accounts: a numbered Method, per-tool
sections with install and six to ten commented invocations, a closing
note on limits, and a worked example from one concrete datum. 14 of 16
OSINT sheets done; the two untouched were already the reference shape.
Flags and endpoints were run or probed rather than recalled, which again
found breakage in what was already published:
- information-laundromat.com is NXDOMAIN; the live tool is
informationlaundromat.com with no hyphen.
- Anonymous Wayback GET /save/<url> returns 429 with a body that looks
like a page, so old scripts fail silently. Replaced with SPN2.
- Bing's Visual Search API was decommissioned in August 2025.
- Public rsshub.app returns 403, and bridge routes fail silently with an
empty feed rather than an error.
- ShadowFinder's CLI is positional under a subcommand now.
- csvkit type inference rewrites 01/03/2026 to 2026-01-03 silently, and
pandas format='mixed' guards against silent NaT coercion.
- wpull is dead and grab-site stale, so browsertrix-crawler is
documented for WARC/WACZ instead.
Also corrects image-video-forensics, which claimed -vsync vfr and
-fps_mode vfr are both accepted: -vsync was removed outright in ffmpeg 9
and fails with 'Unrecognized option'. And trims the /credits ired.team
blurb, which claimed derived sheets in three categories that have none.
The geolocation worked example's two-date-windows claim was checked by
sweeping a full year at five-minute resolution rather than asserted.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Diffstat:
8 files changed, 3030 insertions(+), 225 deletions(-)
diff --git a/src/content/sheets/osint/archiving-and-evidence.md b/src/content/sheets/osint/archiving-and-evidence.md
@@ -4,9 +4,9 @@ description: "Capture sources so they survive deletion, with the hashes and time
category: osint
subcategory: "Archiving & Analysis"
tags: [osint, archiving, evidence, preservation]
-tools: [auto-archiver, wayback, archive-today, yt-dlp]
-difficulty: beginner
-updated: 2026-09-28
+tools: [wget, browsertrix-crawler, yt-dlp, auto-archiver, wayback, archive-today, opentimestamps, openssl, exiftool, shasum]
+difficulty: intermediate
+updated: 2026-10-04
references:
- name: "Bellingcat's Online Investigation Toolkit"
url: "https://bellingcat.gitbook.io/toolkit"
@@ -24,56 +24,547 @@ references:
## What this covers
-Making sure what you found still exists tomorrow, and that you can show it has not changed since
-you found it. Content gets deleted, edited and made private constantly — usually right after
-someone notices attention. Archive first, analyse second.
+Capturing a source so that it still exists after someone deletes it, and so that you can show the
+copy you hold is the copy you took. The capture itself is the easy half; the hash, the third-party
+timestamp and the written provenance are what make it worth anything to an editor, a lawyer or a
+court. Everything below is organised around one question: what does this step actually prove?
## Method
-1. **Archive to something you do not control** — Wayback, archive.today — so the capture has a
- third-party timestamp.
-2. **Also keep a local copy.** Public archives can be removed, rate-limited or unavailable.
-3. **Hash every local file** on acquisition, and record the hash somewhere separate.
-4. **Record provenance:** the URL, the moment you captured it, and how you reached it.
-5. **Capture the media, not a screenshot of it.** A screenshot loses metadata, resolution and
- audio.
+1. **Check what already exists before you capture.** A Wayback snapshot from before the content
+ became interesting is worth more than one you took after. Query the availability and CDX APIs
+ first; if there is an older capture, that is your baseline and the diff against today is itself a
+ finding.
+2. **Capture to something you do not control, and to something you do.** A third-party archive
+ supplies a timestamp nobody can accuse you of forging. A local copy survives the archive being
+ rate-limited, robots-excluded or taken down. Neither substitutes for the other.
+3. **Pick the capture tool by what the page is.** Static HTML and media: `wget --warc-file`.
+ JavaScript-rendered, infinite-scroll or login-walled: a real browser driving the capture, which
+ in practice means `browsertrix-crawler` producing WACZ. Platform video: `yt-dlp`, which gets you
+ the metadata sidecar a screen recording never will.
+4. **Hash on acquisition, before you look at the file.** A hash taken after you cropped, converted
+ or re-saved proves something about your working copy and nothing about the original.
+5. **Timestamp the manifest, not every file.** One hash over the list of hashes, submitted to
+ OpenTimestamps or an RFC 3161 authority, binds the whole collection to a point in time at one
+ cost. That is the step that converts "I have a hash" into "I had this hash on that date".
+6. **Write the provenance down in the same commit as the files.** URL, UTC time of capture, the
+ tool and version, who handed it to you, and how you reached the page. The route to a finding is
+ part of the finding, and it is the part you will forget first.
+7. **Get a second independent capture.** Two archives of the same URL taken by different
+ infrastructure, or the same video from a second uploader, is the difference between a claim and
+ a corroborated claim. Two copies that both trace to one upload are one copy.
+
+Judgement calls worth naming: submitting a URL to a public archive creates a public record that
+somebody is interested in it, so on a target that watches its referrers and its Wayback entries,
+capture locally first and submit later. And a capture of a page you reached while logged in
+contains your session — strip or redact before you share the WARC.
+
+## Key tools
+
+### wget (WARC capture)
+
+The standard way to get a byte-level record of an HTTP exchange rather than a rendered picture of
+it. A WARC stores the request and response headers alongside the body for every resource fetched,
+with a SHA-1 digest per record, so the file is a transcript of the conversation and not just its
+outcome. Use it whenever the page is server-rendered; it is the most portable evidence format in
+this area and every replay tool reads it.
-## Public archives
+```bash
+brew install wget # or: apt install wget
+```
```bash
-# push a URL into the Wayback Machine
-curl -s "https://web.archive.org/save/https://example.com/page"
+# the baseline capture: WARC transcript plus a CDX index of what went into it
+wget --warc-file=capture-001 --warc-cdx \
+ --page-requisites --adjust-extension --no-parent \
+ 'https://example.com/page'
+
+# stamp the operator and case into the warcinfo record, so the file self-documents
+wget --warc-file=capture-002 --warc-cdx \
+ --warc-header="operator: J. Investigator" \
+ --warc-header="description: case-2026-014, post cited in filing" \
+ --warc-header="robots: off" \
+ --page-requisites --adjust-extension 'https://example.com/page'
+
+# leave it uncompressed while you are inspecting it; grep works on a plain WARC
+wget --warc-file=capture-003 --no-warc-compression 'https://example.com/page'
+
+# pull in the CDN that actually serves the images, or the capture renders blank
+wget --warc-file=capture-004 --warc-cdx \
+ --page-requisites --convert-links --adjust-extension \
+ --span-hosts --domains example.com,cdn.example.com \
+ 'https://example.com/page'
+
+# a whole section, politely: the delay is what stops you being blocked mid-capture
+wget --warc-file=capture-005 --warc-cdx --recursive --level=2 --no-parent \
+ --wait=2 --random-wait --limit-rate=500k \
+ -U 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0 Safari/537.36' \
+ 'https://example.com/newsroom/'
+
+# a list of URLs in one WARC, split at 1GB so the files stay movable
+wget --warc-file=capture-006 --warc-max-size=1G --warc-cdx -i urls.txt
-# check existing captures before assuming you need a new one
-curl -s "http://archive.org/wayback/available?url=example.com/page×tamp=20240101" | jq .
+# second pass over the same site without re-storing what you already have
+wget --warc-file=capture-007 --warc-dedup=capture-005.cdx --recursive --no-parent \
+ 'https://example.com/newsroom/'
+
+# read back what the capture contains, without a replay tool
+zcat capture-001.warc.gz | grep -a '^WARC-Target-URI:' | sort -u
+awk 'NR>1 {print $5, $6, $1}' capture-001.cdx # status, SHA-1 digest, URL
```
-`archive.today` (also `archive.ph`, `archive.is`) renders JavaScript-heavy pages that Wayback often
-fails on, and is markedly more resistant to takedown requests. Use both; they fail differently.
+Read the CDX file: the digest column is a base32 SHA-1 of the response payload, which gives you a
+free per-resource integrity check, and the status column tells you which requisites 404'd. A
+capture whose CDX is full of `302` and `403` rows did not get the page.
+
+What this does not prove: the WARC records what *your* client received from *an* IP at that moment.
+It does not prove the page looked that way to anyone else, does not survive geo-targeted or
+personalised content, and carries no third-party attestation of the time — the `WARC-Date` is your
+own clock. `--convert-links` rewrites the mirrored files on disk, not the WARC records, so the
+transcript stays pristine while the browsable copy is usable; keep both. Above all, `wget` executes
+no JavaScript, so on a modern single-page application it faithfully archives an empty shell. If the
+page needs a browser, use one.
+
+### browsertrix-crawler
-## Local capture
+A headless Chrome driven by Webrecorder's crawler, writing WARC and WACZ. This is the live,
+maintained answer for anything JavaScript-rendered, where `wget` returns a loading spinner. It
+scrolls, autoplays, runs site-specific behaviours for the big platforms, and extracts page text for
+search.
+
+The older Python crawlers in this niche have aged out: `wpull` carries a 2013–2016 copyright and no
+release since 2019, and `grab-site` still pins Python 3.7/3.8 and was dropped from nixpkgs after
+23.05. Neither is a reasonable thing to stand behind in 2026. Treat them as read-only history and
+use the crawler below.
```bash
-# single page with everything needed to render it offline
-wget --page-requisites --convert-links --adjust-extension \
- --no-parent --span-hosts --domains example.com,cdn.example.com \
- 'https://example.com/page'
+docker pull webrecorder/browsertrix-crawler
+```
+
+```bash
+# single page, as WACZ, with text extracted for full-text search
+docker run -v "$PWD/crawls:/crawls/" -it webrecorder/browsertrix-crawler crawl \
+ --url 'https://example.com/post/123' --scopeType page \
+ --generateWACZ --text to-pages --collection case-2026-014
+
+# the page plus whatever it links to on the same host, two levels deep
+docker run -v "$PWD/crawls:/crawls/" -it webrecorder/browsertrix-crawler crawl \
+ --url 'https://example.com/newsroom/' --scopeType host --depth 2 \
+ --generateWACZ --collection newsroom
+
+# a feed that only loads on scroll: give the behaviours time to finish
+docker run -v "$PWD/crawls:/crawls/" -it webrecorder/browsertrix-crawler crawl \
+ --url 'https://example.com/feed' --scopeType page-spa \
+ --behaviors autoscroll,autoplay,autofetch,siteSpecific \
+ --behaviorTimeout 300 --generateWACZ --collection feed
+
+# several URLs from a file, one collection, so the manifest covers the set
+docker run -v "$PWD/crawls:/crawls/" -it webrecorder/browsertrix-crawler crawl \
+ --urlFile /crawls/urls.txt --scopeType page --generateWACZ --collection batch-01
+
+# the WACZ is a zip: look inside before you trust it
+unzip -l crawls/collections/case-2026-014/case-2026-014.wacz
+unzip -p crawls/collections/case-2026-014/case-2026-014.wacz datapackage-digest.json | jq .
+
+# the crawl log is JSON Lines — this is where blocked and failed pages surface
+jq -r 'select(.logLevel=="error") | [.timestamp,.message] | @tsv' \
+ crawls/collections/case-2026-014/logs/*.log
+```
+
+A WACZ is a WARC plus an index, a page list and a signed digest manifest, and it opens in
+[ReplayWeb.page](https://replayweb.page/) with no server. That makes it the format to hand to
+someone who is not going to install anything. Check `datapackage-digest.json` after every crawl:
+if the crawler was blocked, you get a valid WACZ containing a block page, and nothing about the
+file itself says so.
+
+What it does not tell you: a browser capture is a recording of one rendering session, which means
+it includes whatever was personalised, A/B-tested or geo-fenced for that session. Running the same
+crawl from a different exit and getting different content is a finding about the site, not an error.
+Logged-in crawls bake the session into the archive — handle accordingly.
+
+### yt-dlp
+
+Platform video and audio with its metadata intact. The point is not the video file; a screen
+recording gets you that. The point is the `.info.json` sidecar, which carries uploader ID, upload
+date, duration, view and comment counts, available formats and often the original title and
+description before anybody edited them.
+
+```bash
+brew install yt-dlp # or: pipx install yt-dlp
+```
+
+```bash
+# the provenance capture: video, metadata, description, thumbnail, subtitles
+yt-dlp --write-info-json --write-description --write-thumbnail \
+ --write-subs --sub-langs 'all' --no-mtime \
+ -o '%(upload_date)s_%(uploader_id)s_%(id)s.%(ext)s' URL
+
+# what formats exist, before you commit to one
+yt-dlp -F URL
+
+# the best single pre-merged file, so you archive bytes the platform served
+yt-dlp -f 'best[ext=mp4]/best' --no-mtime URL
+
+# metadata only — enough to date and attribute a clip without downloading it
+yt-dlp --skip-download --write-info-json --write-thumbnail URL
+
+# comments too; they routinely date an event more precisely than the upload does
+yt-dlp --write-info-json --write-comments --skip-download URL
+
+# a whole channel, newest first, with an archive file so re-runs are incremental
+yt-dlp --write-info-json --no-mtime --download-archive seen.txt \
+ -o '%(upload_date)s_%(id)s.%(ext)s' 'https://example.com/@account/videos'
+
+# the intermediate pages the extractor fetched — useful when an extractor misreads a page
+yt-dlp --write-pages --skip-download URL
+
+# the fields you will actually quote, straight out of the sidecar
+jq -r '[.id,.uploader_id,.upload_date,.timestamp,.duration,.view_count,.title] | @tsv' *.info.json
+```
+
+`--no-mtime` matters: without it the file's mtime is set from the upload date, which silently
+overwrites your own acquisition timeline with the platform's claim. `timestamp` in the sidecar is
+Unix epoch seconds and is usually more precise than `upload_date`, which is date-only and rendered
+in the platform's timezone of choice.
+
+What the sidecar does not establish: every field in it is the platform's assertion, repeated.
+`upload_date` is when it was uploaded, never when it was filmed, and a re-upload resets it.
+Extractors break when platforms change, so a field that is `null` may mean absent or may mean the
+extractor lost it — check against the page. Rate limits and age or region gates will silently
+truncate a channel pull; compare the count you got against the count the channel claims.
+
+### Wayback Machine, from the command line
+
+Three separate endpoints, used at three different moments. Availability answers "is there already a
+capture"; CDX answers "what captures exist, and did the content change between them"; Save Page Now
+creates a new one.
+
+```bash
+# is there a capture at all, and what is the nearest one to a date
+curl -s 'https://archive.org/wayback/available?url=example.com/page×tamp=20240101' | jq .
+
+# every capture of a URL, with the payload digest that reveals real edits
+curl -s 'https://web.archive.org/cdx/search/cdx?url=example.com/page&output=json&fl=timestamp,original,statuscode,digest,length' | jq -r '.[] | @tsv'
+
+# collapse consecutive identical captures: what is left is the change history
+curl -s 'https://web.archive.org/cdx/search/cdx?url=example.com/page&output=json&fl=timestamp,digest&collapse=digest' | jq -r '.[] | @tsv'
+
+# everything ever captured under a host, which is how you find deleted pages
+curl -s 'https://web.archive.org/cdx/search/cdx?url=example.com&matchType=domain&output=json&fl=timestamp,original,statuscode&limit=500' | jq -r '.[] | @tsv'
+
+# only the captures in a window, for a page that mattered on one specific day
+curl -s 'https://web.archive.org/cdx/search/cdx?url=example.com/page&from=20260301&to=20260331&output=json&fl=timestamp,digest' | jq -r '.[] | @tsv'
+
+# submit a new capture: SPN2, authenticated with Internet Archive S3-style keys
+curl -s -X POST -H 'Accept: application/json' \
+ -H "Authorization: LOW $IA_ACCESS_KEY:$IA_SECRET_KEY" \
+ -d 'url=https://example.com/page' -d 'capture_all=1' -d 'capture_screenshot=1' \
+ 'https://web.archive.org/save' | jq .
+
+# poll the job until it reports success, and keep the returned timestamp
+curl -s -H "Authorization: LOW $IA_ACCESS_KEY:$IA_SECRET_KEY" \
+ "https://web.archive.org/save/status/$JOB_ID" | jq '{status, timestamp, original_url}'
+
+# fetch the archived copy itself, and hash what you got
+curl -sL 'https://web.archive.org/web/20260301120000id_/https://example.com/page' | shasum -a 256
+```
+
+Get the key pair from [archive.org/account/s3.php](https://archive.org/account/s3.php) and export
+it as `IA_ACCESS_KEY` / `IA_SECRET_KEY`. **The unauthenticated `GET /save/<url>` form is no longer
+usable for work** — an anonymous request to it returns HTTP 429 rather than a capture. Scripts that
+still rely on it fail silently, because a 429 body looks like a page.
+
+The `id_` infix in that last URL asks for the original bytes without Wayback's navigation banner
+injected, which is what you want if you intend to hash or diff the archived copy. The `digest`
+column in CDX output is the payload hash: two captures sharing a digest are byte-identical, so
+collapsing on it turns a thousand snapshots into the handful of moments the page actually changed.
+That is the single most useful thing the CDX API does.
+
+What Wayback does not give you: permanence. A site can retroactively exclude itself, and captures
+do disappear. It also honours robots directives at crawl time, so absence of a capture is not
+evidence the page did not exist. And a Save Page Now capture is still just a capture — it attests
+that the Internet Archive saw that content at that time, which is a strong third-party claim about
+time and a weak one about authenticity.
+
+### archive.today
+
+A separate archive with a separate failure mode, which is exactly why you use it alongside Wayback.
+It renders pages that Wayback flattens, ignores robots directives, and has a long record of not
+removing things. There is no API and the submission form sits behind bot protection, so this one is
+a browser job — do not script it.
+
+```text
+1. Open https://archive.today (archive.ph and archive.is are the same service;
+ if one mirror is blocked where you are, try another)
+2. Paste the URL into the lower box, "My url is alive and I want to archive its content"
+3. Solve the challenge if you get one. Do not automate this step — it is the step
+ that gets the service to block your address.
+4. Wait for the capture. The result URL looks like https://archive.ph/AbC12
+ and the page header shows the capture time in UTC.
+5. Click "screenshot" in the header to get the full-page render as a separate
+ artefact, and save it. The HTML capture and the screenshot fail differently.
+6. Record the short URL AND the capture time. The short URL is your citation.
+7. Before submitting, use the upper box to search for existing captures of the
+ same URL — the same reason you query Wayback's CDX API first.
+```
+
+Two things to know. The service is deliberately opaque about its infrastructure and funding, which
+means you should not treat it as your only copy of anything; keep the local WARC. And it fetches
+the page itself, from its own address, so a submission does not leak your IP to the target — but it
+does create a public, searchable record that someone archived that URL at that minute.
+
+### Hashing and the manifest
+
+Hashes are the whole evidentiary argument. A hash recorded at acquisition and published or
+timestamped separately lets you show, later, that the file you are producing is the file you took.
+The manifest pattern — one file listing every hash, then one hash of the manifest — scales that to a
+collection without timestamping a thousand files.
+
+```bash
+# macOS ships shasum; coreutils ships sha256sum. Same output format.
+shasum -a 256 evidence.mp4
+sha256sum evidence.mp4
+```
+
+```bash
+# hash the whole capture tree in a stable order, so the manifest is reproducible
+find ./capture -type f -print0 | sort -z | xargs -0 shasum -a 256 > MANIFEST.sha256
+
+# the manifest's own hash: this single line is what you timestamp and publish
+shasum -a 256 MANIFEST.sha256 | tee MANIFEST.sha256.sha256
+
+# verify the collection, any time later
+shasum -a 256 -c MANIFEST.sha256
+
+# show only what failed, which is what you actually want on a large set
+shasum -a 256 -c MANIFEST.sha256 2>/dev/null | grep -v ': OK$'
+
+# a provenance record that travels with the files, in the same directory
+cat > PROVENANCE.txt <<'EOF'
+case: 2026-014
+url: https://example.com/post/123
+captured_at: 2026-10-04T04:07:56Z # UTC, from `date -u +%FT%TZ`
+captured_by: J. Investigator
+tooling: wget 1.25.0 --warc-file; yt-dlp 2026.08.19
+route: linked from https://example.com/newsroom/ on 2026-10-03
+notes: page required no login; no personalisation observed
+EOF
+
+# fold the provenance into the manifest so the two cannot drift apart
+shasum -a 256 PROVENANCE.txt >> MANIFEST.sha256
+
+# find duplicates across two collections by hash, not by filename
+sort MANIFEST.sha256 other/MANIFEST.sha256 | awk '{print $1}' | sort | uniq -d
+```
+
+Use SHA-256. MD5 and SHA-1 still appear in forensic tooling and are fine as identifiers, but a
+collision-prone hash invites an argument you do not need to have. Note that `find | sort` matters:
+an unsorted manifest changes order between runs and filesystems, which makes two manifests of the
+same files look different.
+
+What a hash proves and does not: it proves the bytes have not changed since the hash was taken. It
+says nothing about when they were taken, nothing about where they came from, and nothing about
+whether the content is true. On its own a hash in your own notes is worth very little, because you
+could have written it at any time — which is what the next section fixes.
+
+### Timestamping: OpenTimestamps and RFC 3161
+
+The step that turns a hash into evidence. A timestamp is a third party's signed assertion that a
+given digest existed before a given moment. Without it, your hash only proves internal
+consistency; with it, you can show you held that exact file on that date and could not have
+produced it later.
+
+```bash
+pipx install opentimestamps-client # the `ots` command
+# openssl is already present for the RFC 3161 route
+```
+
+```bash
+# OpenTimestamps: anchors the digest in the Bitcoin blockchain, free, no account
+ots stamp MANIFEST.sha256
+
+# what the proof currently contains: pending calendar attestations, then a block
+ots info MANIFEST.sha256.ots
+
+# calendars need an hour or so to get into a block; upgrade the proof afterwards
+ots upgrade MANIFEST.sha256.ots
+
+# verify. Full verification wants a local Bitcoin node; without one you are
+# trusting the calendar, which is weaker but still a third party
+ots verify MANIFEST.sha256.ots
-# video and audio, with metadata and subtitles preserved
-yt-dlp --write-info-json --write-subs --write-thumbnail \
- --no-mtime -o '%(upload_date)s-%(id)s.%(ext)s' URL
+# verify a digest you were given, without holding the file
+ots verify -d 355c357708e4840d12b0d4284d9fe1a911c79c013953874f03068b81127f3363 MANIFEST.sha256.ots
-# hash everything on acquisition
-find ./capture -type f -exec sha256sum {} \; | tee capture.sha256
+# RFC 3161 route: build a timestamp query over the file's SHA-256
+openssl ts -query -data MANIFEST.sha256 -sha256 -cert -out manifest.tsq
-# verify later
-sha256sum -c capture.sha256
+# send it to a Time Stamping Authority and keep the signed reply
+curl -s -H 'Content-Type: application/timestamp-query' \
+ --data-binary @manifest.tsq https://freetsa.org/tsr -o manifest.tsr
+
+# read the token: the TSA's time, its identity and the digest it signed over
+openssl ts -reply -in manifest.tsr -text
+
+# verify the token against the TSA's chain — this is the check a reviewer repeats
+curl -sO https://freetsa.org/files/cacert.pem
+curl -sO https://freetsa.org/files/tsa.crt
+openssl ts -verify -data MANIFEST.sha256 -in manifest.tsr \
+ -CAfile cacert.pem -untrusted tsa.crt
+```
+
+A successful verification prints `Verification: OK`, and that line is the thing you cite. Keep the
+`.tsr`, the CA chain you verified against, and the exact file — the token is signed over the digest,
+so a reviewer needs the file to recompute it.
+
+Which route to use: OpenTimestamps costs nothing, needs no account, and anchors to a public
+blockchain, so the proof outlives any single company — but it is coarse (block granularity, roughly
+an hour) and full verification wants a Bitcoin node. An RFC 3161 TSA gives you a precise, signed,
+PKI-rooted token that institutional reviewers already recognise, at the price of trusting that TSA
+and its certificate remaining validatable. Do both; they are cheap and they fail differently.
+
+What a timestamp does not prove: that the content is authentic, that the capture was complete, or
+that you obtained it lawfully. It proves the file existed no later than that moment. That is a
+narrow claim, and stating it narrowly is what makes it credible.
+
+### Bellingcat Auto Archiver
+
+The pipeline to set up when you are collecting continuously rather than case by case. It reads URLs
+from a feeder, runs a platform-specific extractor, then runs enrichers that hash, thumbnail,
+timestamp and push to Wayback, and writes the results to a database you can read as a case index.
+It is the only tool here that performs the hash-and-timestamp discipline on every item without you
+remembering to.
+
+```bash
+pipx install auto-archiver
+auto-archiver --help # the flag list is generated from the installed modules
+```
+
+Configuration is a single YAML file; the default path is `secrets/orchestration.yaml`. Only modules
+named under `steps` are loaded, so the file is both configuration and a declaration of the pipeline.
+
+```yaml
+# orchestration.yaml
+steps:
+ feeders: [cli_feeder]
+ extractors: [generic_extractor]
+ enrichers: [hash_enricher, meta_enricher, thumbnail_enricher,
+ timestamping_enricher, opentimestamps_enricher]
+ databases: [csv_db, console_db]
+ storages: [local_storage]
+ formatters: [html_formatter]
+
+hash_enricher:
+ algorithm: SHA-256 # the only alternative is SHA3-512
+
+timestamping_enricher:
+ tsa_urls:
+ - http://timestamp.identrust.com
+ - http://zeitstempel.dfn.de
+ allow_selfsigned: false # leaving this false is the whole point of the module
+
+local_storage:
+ save_to: ./archived
+ path_generator: url # flat | url | random
+ filename_generator: static # names the file after its hash
+
+csv_db:
+ csv_file: ./archived/results.csv
+
+logging:
+ level: INFO
+ file: ./archived/archiver.log
+```
+
+```bash
+# one URL through the whole pipeline
+auto-archiver --config orchestration.yaml 'https://example.com/post/123'
+
+# several at once; the CSV gets one row per URL
+auto-archiver --config orchestration.yaml 'https://example.com/a' 'https://example.com/b'
+
+# no config file at all, for a one-off capture with sane defaults
+auto-archiver --mode simple 'https://example.com/post/123'
+
+# override the enricher list without editing the file
+auto-archiver --config orchestration.yaml \
+ --enrichers hash_enricher timestamping_enricher opentimestamps_enricher \
+ 'https://example.com/post/123'
+
+# push to Wayback as part of the run, with the same Internet Archive key pair
+auto-archiver --config orchestration.yaml \
+ --enrichers hash_enricher wayback_extractor_enricher \
+ --wayback_extractor_enricher.key "$IA_ACCESS_KEY" \
+ --wayback_extractor_enricher.secret "$IA_SECRET_KEY" \
+ 'https://example.com/post/123'
+
+# feed a column of URLs from a spreadsheet export
+auto-archiver --config orchestration.yaml \
+ --feeders csv_feeder --csv_feeder.files urls.csv --csv_feeder.column link
+
+# turn up logging when an extractor silently returns nothing
+auto-archiver --config orchestration.yaml --logging.level DEBUG 'https://example.com/post/123'
+
+# the results CSV is the case index: one row per URL, with hashes and archive links
+column -s, -t < ./archived/results.csv | less -S
+```
+
+Two enrichers do the evidentiary work and they are not interchangeable. `timestamping_enricher`
+aggregates the item's file hashes into one text file and gets an RFC 3161 token over it from each
+TSA in the list; `opentimestamps_enricher` anchors the same digest via the OpenTimestamps
+calendars. Note that `freetsa.org` is deliberately absent from the shipped TSA defaults because its
+certificate does not validate against system authorities — you can add it, but only together with
+`allow_selfsigned: true`, which weakens exactly the property you wanted.
+
+The `gsheet_feeder_db` module is what makes this a team tool: analysts paste URLs into a Google
+Sheet, the archiver writes the hash, the stored path and the archive links back into adjacent
+columns. The `wacz_extractor_enricher` shells out to Docker for browser-based capture, and
+`wayback_extractor_enricher` needs the same key pair as Save Page Now.
+
+Watch the version. The client checks on startup and says so when it is behind, and per-platform
+extractors break often enough that an old release produces empty captures that look like
+successes. After every run, scan the results CSV for rows that carry a hash but no media.
+
+### ExifTool, for capture-time metadata
+
+Covered in depth on [Image & Video Forensics](/sheets/osint/image-video-forensics); the archiving
+use is narrower. You are not asking whether the file is manipulated, you are recording what the
+file asserted about itself at the moment it entered your custody, so that a later disagreement is
+about the metadata rather than about your handling.
+
+```bash
+# the full tag dump, archived next to the hash and never edited again
+exiftool -j -g1 -a -u capture/evidence.mp4 > capture/evidence.mp4.exif.json
+
+# the fields that date and attribute a capture, in one line
+exiftool -CreateDate -ModifyDate -FileModifyDate -GPSPosition -Make -Model -Software \
+ capture/evidence.mp4
+
+# every timestamp in the file, grouped by where it came from
+exiftool -time:all -a -G1 -s capture/evidence.mp4
+
+# the whole capture tree as one CSV, which doubles as a collection inventory
+exiftool -r -csv -FileName -FileSize -MIMEType -CreateDate -GPSPosition ./capture/ > inventory.csv
+
+# flag files whose structure does not match their declared format
+exiftool -r -validate -warning -a ./capture/
+
+# strip metadata from a working copy before you publish, to protect a source
+cp capture/evidence.jpg publish/evidence.jpg
+exiftool -all= -overwrite_original publish/evidence.jpg
+shasum -a 256 publish/evidence.jpg # a different file, so a different hash: say so
```
-`--no-mtime` matters: without it `yt-dlp` sets the file's modification time from the upload date,
-which muddles your own acquisition timeline.
+Run this on the copy, after hashing the original, and never with `-overwrite_original` on anything
+in the capture tree. Note the last block: a published, stripped file is a *different artefact* with a
+different hash, and conflating the two is a straightforward way to be accused of altering evidence.
+Record both hashes and the relationship between them.
-## Purpose-built tooling
+What it does not tell you: metadata is written as easily as it is read, so a helpful
+`CreateDate` is consistent-with and never proof. Most files pulled from platforms have had their
+metadata stripped on upload, and that absence is normal rather than suspicious.
+
+## Case-management and pipeline tools
| Tool | What it does |
| --- | --- |
@@ -83,14 +574,6 @@ which muddles your own acquisition timeline.
| [Lumen](https://lumendatabase.org/) | Archive of takedown notices — sometimes the only record that something existed. |
| [Web Archives](https://github.com/dessant/web-archives) | Browser extension that queries many archive services at once. |
-Auto Archiver is the one to set up if you are collecting on an ongoing basis rather than
-case-by-case:
-
-```bash
-pipx install auto-archiver
-auto-archiver --config orchestration.yaml
-```
-
## Tool reference
| Tool | What it does | Cost |
@@ -111,6 +594,119 @@ auto-archiver --config orchestration.yaml
nothing about the original.
- **Archiving can notify.** Submitting a URL to a public archive creates a public record that
someone is interested in it.
+- **A capture that succeeded is not a capture that worked.** A WARC full of 403s, a WACZ containing
+ a block page, and a `yt-dlp` channel pull truncated by rate limiting all exit zero. Read the CDX
+ status column, the crawl log and the row count before you file anything.
+- **Old scripts hitting `GET /save/<url>` now get HTTP 429.** Anonymous Save Page Now is no longer
+ usable; a 429 body is still a body, so a script that does not check the status silently records a
+ failure as a success. Use the authenticated SPN2 POST.
+- **A hash in your own notes is nearly worthless.** You could have written it at any time. The
+ third-party timestamp is what makes it an assertion about the past rather than about your memory.
+
+## Worked example
+
+One datum: a URL to a post carrying a video clip, `https://example.com/@account/post/123`, cited in
+a filing you have been asked to check. The URL, the hashes and the timestamps below are invented;
+the order of the steps and what each one proves are not.
+
+```bash
+# 1. what already exists, before you touch the page
+curl -s 'https://web.archive.org/cdx/search/cdx?url=example.com/@account/post/123&output=json&fl=timestamp,digest,statuscode&collapse=digest' | jq -r '.[] | @tsv'
+# timestamp digest statuscode
+# 20260228183044 H4XQ... 200
+# 20260302094112 ZK7B... 200
+```
+
+Two distinct payload digests, two days apart. Collapsing on digest is what revealed that: the page
+was edited between 28 February and 2 March, which is a finding in itself and the reason to pull
+both old captures before capturing the live page.
+
+```bash
+# 2. local capture of the live page, byte-level, with the case in the warcinfo record
+mkdir -p case-2026-014 && cd case-2026-014
+wget --warc-file=live --warc-cdx --no-warc-compression \
+ --warc-header="operator: J. Investigator" \
+ --warc-header="description: case-2026-014, post cited at para 17" \
+ --page-requisites --adjust-extension --no-parent \
+ 'https://example.com/@account/post/123'
+awk 'NR>1 {print $5, $1}' live.cdx | sort | uniq -c
+# 14 200 https://example.com/...
+# 3 403 https://cdn.example.com/media/...
+```
+
+Three requisites returned 403, and they are the media files. The `wget` capture has the page text
+and not the clip, which is the common outcome and the reason the next two steps exist rather than
+being optional extras.
+
+```bash
+# 3. the rendered page, in a browser, because the post body loads via JavaScript
+docker run -v "$PWD/crawls:/crawls/" -it webrecorder/browsertrix-crawler crawl \
+ --url 'https://example.com/@account/post/123' --scopeType page \
+ --generateWACZ --text to-pages --collection case-2026-014
+unzip -p crawls/collections/case-2026-014/case-2026-014.wacz datapackage-digest.json | jq .
+# { "path": "datapackage.json", "hash": "sha256:9f2c...", "signedData": null }
+```
+
+`signedData: null` means the WACZ is unsigned — fine, because you are about to timestamp it
+yourself. The digest over `datapackage.json` is the crawler's own integrity check on the archive;
+record it, then stop relying on it and use your own hash.
+
+```bash
+# 4. the media, with its provenance sidecar, which the WARC never got
+yt-dlp --write-info-json --write-description --write-thumbnail --no-mtime \
+ -o '%(upload_date)s_%(uploader_id)s_%(id)s.%(ext)s' \
+ 'https://example.com/@account/post/123'
+jq -r '[.id,.uploader_id,.upload_date,.timestamp,.duration,.view_count] | @tsv' *.info.json
+# 123 account 20260226 1772120000 47 18422
+```
+
+The sidecar says the clip was uploaded on 26 February — two days *before* the earliest Wayback
+capture of the post, and before the edit the CDX diff exposed. That is the pivot: the video predates
+the post text it now sits under, so the claim to check is whether the caption was changed, not
+whether the footage is real.
+
+```bash
+# 5. hash everything, in a stable order, before any further handling
+cd .. && find ./case-2026-014 -type f -print0 | sort -z | xargs -0 shasum -a 256 > MANIFEST.sha256
+wc -l MANIFEST.sha256 && shasum -a 256 MANIFEST.sha256
+# 31 MANIFEST.sha256
+# 4b1c8e... MANIFEST.sha256
+```
+
+```bash
+# 6. the step that makes the hash mean something: two independent timestamps
+ots stamp MANIFEST.sha256
+openssl ts -query -data MANIFEST.sha256 -sha256 -cert -out manifest.tsq
+curl -s -H 'Content-Type: application/timestamp-query' \
+ --data-binary @manifest.tsq https://freetsa.org/tsr -o manifest.tsr
+openssl ts -reply -in manifest.tsr -text | grep -E 'Status:|Time stamp:'
+# Status: Granted.
+# Time stamp: Oct 4 04:06:07 2026 GMT
+```
+
+Now the claim is narrow and defensible: this set of 31 files, with these hashes, existed no later
+than 04:06 UTC on 4 October 2026, attested by a TSA and by the Bitcoin blockchain, neither of which
+you control.
+
+```bash
+# 7. a third-party capture, submitted last so the local copy exists first
+curl -s -X POST -H 'Accept: application/json' \
+ -H "Authorization: LOW $IA_ACCESS_KEY:$IA_SECRET_KEY" \
+ -d 'url=https://example.com/@account/post/123' -d 'capture_all=1' \
+ 'https://web.archive.org/save' | jq -r '.job_id'
+# spn2-9f3c...
+```
+
+Then the same URL by hand into [archive.today](https://archive.today), and its short URL and UTC
+capture time into `PROVENANCE.txt`. Submitting last is deliberate: both submissions are public, so
+if the account notices and deletes the post, you already hold the capture.
+
+What this run establishes: the page as it stood at a timestamped moment, the clip with the
+platform's own metadata, and a change history from Wayback showing the post was edited after the
+video was uploaded. What it does not establish: that the footage shows what the caption claims, or
+where or when it was filmed. Those are
+[Image & Video Forensics](/sheets/osint/image-video-forensics) and
+[Geolocation](/sheets/osint/geolocation) questions, and no amount of archiving substitutes for them.
## Broader catalogues
diff --git a/src/content/sheets/osint/data-analysis-and-visualisation.md b/src/content/sheets/osint/data-analysis-and-visualisation.md
@@ -4,9 +4,9 @@ description: "Clean messy collected data, map relationships between entities, bu
category: osint
subcategory: "Archiving & Analysis"
tags: [osint, analysis, visualisation, graphs, timelines]
-tools: [openrefine, gephi, datasette, datawrapper, qgis]
+tools: [openrefine, csvkit, sqlite-utils, datasette, pandas, networkx, gephi, datawrapper, pinpoint]
difficulty: intermediate
-updated: 2026-09-28
+updated: 2026-10-04
references:
- name: "Bellingcat's Online Investigation Toolkit"
url: "https://bellingcat.gitbook.io/toolkit"
@@ -24,81 +24,471 @@ references:
## What this covers
-The part after collection. Thousands of rows of scraped posts, a list of company officers, a set of
-geotagged images — none of it means anything until it is cleaned, related and shown. Analysis is
-also where most errors enter an investigation, because a chart makes a weak claim look strong.
+The part after collection, where a pile of scraped rows becomes a claim you are willing to publish.
+Cleaning, relating, sequencing and showing are four separate jobs with four different failure
+modes, and analysis is where most errors enter an investigation because a chart makes a weak claim
+look strong. Every step below either loses rows or changes their meaning, so the discipline is to
+know which, and to write it down.
+
+## Method
+
+1. **Count the rows you started with, and count them again after every step.** A step that drops
+ rows silently is the single most common way a dataset ends up saying the wrong thing. Record the
+ number before and after; if it changed and you did not intend it, stop.
+2. **Normalise before you deduplicate, deduplicate before you count.** `Northgate Ltd`,
+ `Northgate Limited` and `northgate ltd` are three rows and one company. Any count you compute
+ before merging them is wrong, and no later step fixes it.
+3. **Disable type inference on first contact.** Tools guess, and a date guessed month-first turns
+ 1 March into 3 January without a warning. Load everything as text, inspect, then convert
+ deliberately with a format you specified.
+4. **Decide what the unit of analysis is before you join.** One row per post, per account, per
+ account-day? A join that silently fans out one-to-many inflates every count downstream, and the
+ resulting chart looks fine.
+5. **Keep raw, cleaned and derived as three separate files.** The raw file is never edited. The
+ cleaning is a script, not a sequence of clicks you cannot repeat. The derived file is what the
+ chart reads.
+6. **Mark the gaps.** A collection outage renders as a quiet period and reads as an event. If you
+ know when your collection was incomplete, that belongs on the timeline as a shaded band, not as
+ a footnote nobody reads.
+7. **Choose the figure that makes the weakest honest claim.** If the finding is "these two accounts
+ post within sixty seconds of each other, 94 times", say that. A force-directed graph of the whole
+ dataset with those two nodes somewhere in it is prettier and claims more than you can support.
+
+## Key tools
+
+### OpenRefine
+
+The cleaning tool to learn properly, because its clustering is the one feature that no CLI
+replicates well: it proposes groups of values that are probably the same thing and lets you merge
+each group in one click, with the count of affected rows visible as you decide. That visibility is
+the point — merging is a judgement, and OpenRefine makes you make it row-count in hand.
+
+Download the current release (3.10.0, February 2026) from
+[openrefine.org/download](https://openrefine.org/download) and run it locally; it serves a browser
+UI on `127.0.0.1:3333` and your data never leaves the machine.
+
+Clustering lives under a column's dropdown, `Edit cells` → `Cluster and edit`. Work the methods in
+this order:
+
+```text
+key collision / fingerprint lowercases, strips punctuation and diacritics, sorts
+ the words. Start here: almost no false positives.
+ Catches "SMITH, John" = "John Smith".
+key collision / ngram-fingerprint n=2 catches typos and missing spaces that fingerprint
+ misses. Raise n to loosen, lower it to tighten.
+key collision / metaphone3 English phonetics: "Stephen" = "Steven". Use
+ cologne-phonetic for German, daitch-mokotoff for
+ Slavic and Yiddish names, beider-morse for a stricter pass.
+nearest neighbour / levenshtein edit distance, with a radius you set. Slow. This is
+ where real false positives start, so read every group.
+nearest neighbour / ppm compression-based, for long strings. Last resort; it
+ over-merges short values badly.
+```
+
+GREL, in the `Edit cells` → `Transform` box, is for everything clustering cannot do:
+
+```text
+value.trim() strip the whitespace that breaks every join
+value.replace(/\s+/, " ").trim() collapse runs of whitespace to one space
+fingerprint(value) the clustering key itself, as a new column,
+ so you can group on it in SQL later
+value.replace(/\b(Ltd|Limited|PLC|LLC|Inc)\b\.?/i, "").trim()
+ drop corporate suffixes before matching names
+value.toDate("dd/MM/yyyy") parse an unambiguous day-first date explicitly
+value.toDate(false, "dd/MM/yyyy", "yyyy-MM-dd") try day-first, then ISO; false = not month-first
+value.toDate("dd/MM/yyyy").datePart("hours") pull a component back out
+diff(value.toDate("yyyy-MM-dd"), cells["first_seen"].value.toDate("yyyy-MM-dd"), "days")
+ day gap between two columns
+value.match(/^(\+?\d{1,3})[\s-]?(\d{6,})$/)[1] capture group from a regex, as a value
+value.splitByCharType().join("|") expose where letters and digits meet, which
+ is how you find "acct12" vs "acct 12"
+cross(value, "officers", "company_name")[0].cells["role"].value
+ look a value up in another OpenRefine project
+```
+
+Reconciliation is the other thing worth the setup: it matches a column of names against an external
+authority and returns scored candidates rather than a guess. Add a service under
+`Reconcile` → `Start reconciling` → `Add standard service`:
+
+```text
+https://wikidata.reconci.link/en/api Wikidata; swap "en" for another language code
+```
+
+That host now 307-redirects to `wikidata-reconciliation.wmcloud.org`, which is the current canonical
+home — either URL works, the second avoids the hop. The older
+`wdreconcile.toolforge.org` service is deprecated and unreliable; do not point new projects at it.
+
+Read reconciliation output as candidates, never as answers. The service returns a match score and
+OpenRefine auto-matches above a threshold, which means it will confidently bind a small company to
+a famous one of a similar name. Set the threshold high, review the auto-matches, and keep the
+unmatched rows visible rather than dropping them — an unmatched name is information about your
+authority's coverage, not about the entity.
+
+### csvkit
+
+The right tool between "I have a CSV" and "I have a database": it does the inspection, cutting,
+filtering and joining that would otherwise be a throwaway script, and it streams, so it works on
+files larger than memory. Reach for it before pandas when the job is shaped like a shell pipeline.
+
+```bash
+pipx install csvkit
+```
+
+```bash
+# what are the columns actually called, and in what order
+csvcut -n collected.csv
-## Cleaning
+# does the file parse at all: ragged rows, empty columns, mismatched header
+csvclean -a --label - collected.csv > /dev/null
-Collected data is always dirty: inconsistent name spellings, mixed date formats, duplicate entities
-under slightly different labels.
+# a readable look at the first rows, before you decide anything
+csvcut -c 1-6 collected.csv | csvlook --max-rows 20
-**OpenRefine** is the right tool and is underused. Its clustering function finds values that are
-probably the same thing — `Jon Smith`, `John Smith`, `SMITH, John` — and lets you merge them in one
-pass. Do this before any counting, or your counts are wrong.
+# the whole statistical profile: type, nulls, unique count, min, max, frequencies
+csvstat collected.csv
+
+# just the questions that matter on first contact
+csvstat --nulls collected.csv
+csvstat -c company --freq --freq-count 25 collected.csv
+
+# convert from whatever you were given, naming the sheet explicitly
+in2csv -f xlsx --sheet 'Payments' register.xlsx > payments.csv
+in2csv -f json -k results api_dump.json > api.csv
+
+# filter by regex, case-insensitively, keeping the header
+csvgrep -c company -r '(?i)northgate' collected.csv > northgate.csv
+
+# rows where ANY column matches, for a name you cannot locate in a known column
+csvgrep -a -r '(?i)northgate' collected.csv
+
+# join two collections on an id, keeping every left row so you can see the misses
+csvjoin -I -c id --left payments.csv officers.csv > merged.csv
+
+# stack files from separate collection runs, tagging each with its source
+csvstack -g 2026-02,2026-03 -n batch feb.csv mar.csv > all.csv
+
+# ad-hoc SQL over CSVs without creating a database
+csvsql --query "SELECT company, COUNT(*) n, SUM(amount) total
+ FROM payments GROUP BY 1 HAVING n > 1 ORDER BY total DESC" payments.csv
+```
+
+**Use `-I` on first contact, every time.** csvkit infers types by default, and the inference is
+month-first: a column containing `01/03/2026` comes out of a plain `csvjoin` as `2026-01-03`, so
+1 March silently becomes 3 January, with no warning and no error. `-I` passes the value through
+verbatim. If you do want parsing, say what the format is with `--date-format '%d/%m/%Y'` — and note
+that a column with genuinely mixed formats will defeat that too, which is a reason to split the
+column rather than to trust the parser.
+
+Two other limits worth knowing: `csvjoin` reads every input fully into memory and says so in its
+own help, so a multi-gigabyte join belongs in SQLite rather than here; and the `--blanks` flag
+exists because csvkit converts `""`, `na`, `n/a`, `none` and `.` to NULL by default, which is
+usually right and is occasionally the destruction of a meaningful category.
+
+### sqlite-utils
+
+The step that turns a directory of CSVs into something you can actually interrogate. It builds the
+schema from the data, handles the inserting, and gives you full-text search and foreign keys
+without writing DDL. Past roughly a hundred thousand rows this is where analysis should live rather
+than in a dataframe.
+
+```bash
+pipx install sqlite-utils
+```
+
+```bash
+# CSV straight into a table, with a primary key so re-runs are idempotent
+sqlite-utils insert case.db payments payments.csv --csv --pk id
+
+# everything as TEXT, which is the safe default when the dates are ambiguous
+sqlite-utils insert case.db payments_raw payments.csv --csv --no-detect-types
+
+# newline-delimited JSON, e.g. a scraper's output, flattening nested objects
+sqlite-utils insert case.db posts posts.jsonl --nl --flatten --pk id
+
+# re-run a collection without duplicating: replace rows that already exist
+sqlite-utils insert case.db posts new.jsonl --nl --pk id --replace
+
+# add a column from the second batch without rebuilding the table
+sqlite-utils insert case.db posts batch2.jsonl --nl --pk id --alter
+
+# inspect what you built
+sqlite-utils schema case.db
+sqlite-utils tables case.db --counts
+
+# full-text search across the columns that carry names and free text
+sqlite-utils enable-fts case.db payments name company --create-triggers
+sqlite-utils search case.db payments 'northgate' -c id -c company
+
+# the normalisation that makes counting honest: one row per distinct company
+sqlite-utils extract case.db payments company --table companies --fk-column company_id
+
+# a query, as JSON, straight out of the CLI
+sqlite-utils case.db "SELECT company, COUNT(*) n FROM payments GROUP BY 1 ORDER BY n DESC LIMIT 20"
+
+# throwaway query against CSVs with no database at all
+sqlite-utils memory payments.csv officers.csv \
+ "SELECT p.company, o.role FROM payments p JOIN officers o ON p.id = o.id"
+
+# indexes, once a query starts being slow rather than instant
+sqlite-utils create-index case.db payments company date
+```
+
+`--pk` is the one flag to never omit: without it every re-run of a collection appends duplicates,
+and you will not notice until a count is wrong. Note that type detection is on by default — the
+flag is `--no-detect-types` to turn it off, and there is no `--detect-types`. `extract` is the
+underused one: it pulls a repeated text column into its own table with a foreign key, which is both
+the normalisation step and a cheap way to see how many distinct entities you really have.
+
+What it does not do: validate anything. `sqlite-utils` will happily build a table where `amount` is
+`TEXT` because one row contained `n/a`, and every `SUM` over it then returns zero without
+complaining. Check the schema after every insert.
+
+### Datasette
+
+Publishing a dataset other people can query, without building an application. Point it at the
+SQLite file and you get faceted browsing, arbitrary SQL in the URL, CSV and JSON export, and a
+permalink per row — which means a colleague can cite a specific record and you can both see the
+query that produced a number.
+
+```bash
+pipx install datasette
+```
+
+```bash
+# serve locally, read-only, and open a browser
+datasette serve -i case.db -o
+
+# bind somewhere a colleague on the network can reach, with a port you chose
+datasette serve -i case.db -h 0.0.0.0 -p 8080
+
+# several databases at once, with cross-database joins enabled
+datasette serve -i case.db -i reference.db --crossdb
+
+# source, licence and column descriptions, which is the difference between a
+# dataset and a pile of rows
+datasette serve -i case.db -m metadata.yml
+
+# ad-hoc: query CSVs in memory without building a database first
+datasette --memory
+
+# a scripted query against the JSON API, for a figure you will regenerate
+datasette serve -i case.db --get '/case.json?sql=select+company,count(*)+from+payments+group+by+1'
+
+# publish to a container host
+datasette publish cloudrun case.db --service case-2026-014
+datasette publish heroku case.db --name case-2026-014
+```
-For anything scriptable, pandas:
+`-i` opens the file immutable, which is what you want for published evidence: nothing served can
+alter the database, and Datasette can cache aggressively because the contents cannot change.
+
+Note that `datasette publish` ships only `cloudrun` and `heroku` targets. Vercel, Fly and Cloud Run
+with custom settings are plugins (`datasette-publish-vercel`, `datasette-publish-fly`) that must be
+installed first — a `datasette publish vercel` invocation fails with an unknown-command error on a
+bare install. Current stable is the 0.6x line; the 1.0 alphas change the plugin and metadata APIs,
+so pin your version if you are publishing something you need to rebuild identically later.
+
+The caution with Datasette is social rather than technical: a published database is far more
+exposing than a chart. Every row is readable, every join is possible, and arbitrary SQL means
+anybody can compute something you did not anticipate. Before publishing, drop the columns you only
+needed for cleaning, and remember that a `source_url` column can re-identify people a redacted name
+column was meant to protect.
+
+### pandas, for timelines
+
+Sequence is the most load-bearing and most abused structure in open-source work: who posted first,
+how fast something propagated, whether an account's activity pattern changed. Build it in pandas
+because the operations you need — timezone conversion, resampling, gap measurement — are one line
+each and are exactly the ones a spreadsheet gets wrong.
+
+```bash
+pipx install pandas
+```
```python
import pandas as pd
-df = pd.read_csv("collected.csv")
-df["date"] = pd.to_datetime(df["date"], errors="coerce", utc=True)
-df["name"] = df["name"].str.strip().str.casefold()
-df = df.drop_duplicates(subset=["name", "date"])
+df = pd.read_csv("collected.csv", dtype=str) # everything as text; convert deliberately
-# what did parsing fail on? this is where silent data loss hides
-print(df["date"].isna().sum(), "unparseable dates")
-```
+# parse explicitly, and keep the failures visible rather than dropping them
+df["ts"] = pd.to_datetime(df["ts"], errors="coerce", utc=True, format="mixed")
+bad = df["ts"].isna()
+print(f"{bad.sum()} of {len(df)} timestamps unparseable")
+df.loc[bad, ["id", "ts"]].to_csv("unparsed_timestamps.csv", index=False)
+df = df.dropna(subset=["ts"]).sort_values("ts")
-Always check what failed to parse. Rows silently dropped by a coercion are the classic way a
-dataset ends up telling you the wrong thing.
+# volume per day, with empty days present as zero rather than missing
+print(df.set_index("ts").resample("1D").size())
-## Relationships
+# posting hours in the timezone the actor plausibly lives in, not in UTC
+local = df["ts"].dt.tz_convert("Europe/Kyiv")
+print(local.dt.hour.value_counts().sort_index())
-**Gephi** for network graphs — people, companies, accounts and the edges between them. The useful
-outputs are usually degree (who is most connected), betweenness (who bridges otherwise separate
-clusters) and modularity (what the communities are). Force-directed layouts look impressive and say
-little on their own; the metrics are the finding.
+# activity heatmap: hour of day by actor, which is how a shift pattern shows up
+print(df.groupby([local.dt.hour, "actor"]).size().unstack(fill_value=0))
-**Maltego** automates collection *into* a graph via transforms, which is convenient but ties you to
-its data sources.
+# gaps between consecutive events, in seconds — the coordination signal
+df["gap_s"] = df["ts"].diff().dt.total_seconds()
+print(df.nsmallest(20, "gap_s").loc[:, ["ts", "actor", "gap_s"]])
-For a company-ownership chain, a graph is usually overkill — a simple parent/subsidiary tree is
-clearer and harder to misread.
+# near-simultaneous posts by different actors, which is the claim worth making
+cols = ["ts", "actor", "id"]
+pairs = pd.merge_asof(
+ df.loc[:, cols].rename(columns={"actor": "a", "id": "id_a"}),
+ df.loc[:, cols].rename(columns={"actor": "b", "id": "id_b"}),
+ on="ts", tolerance=pd.Timedelta("60s"), direction="nearest", allow_exact_matches=False,
+)
+print(pairs.query("a != b").groupby(["a", "b"]).size().sort_values(ascending=False).head(20))
-## Querying
+# known collection gaps, recorded as data so they reach the chart
+gaps = pd.DataFrame({"start": pd.to_datetime(["2026-03-05T00:00Z"]),
+ "end": pd.to_datetime(["2026-03-07T12:00Z"])})
+gaps.to_csv("collection_gaps.csv", index=False)
-**Datasette** turns a SQLite file into a browsable, queryable, publishable web interface. It is the
-fastest route from "I have a CSV" to "colleagues can explore this".
+df.to_csv("timeline.csv", index=False) # the derived file the chart reads
+```
+
+**`format="mixed"` is not optional.** Without it, pandas infers a single format from the first
+value and silently coerces every differently-formatted-but-perfectly-valid value to `NaT`: a column
+holding `2026-03-01T08:14:00Z` and `2026-03-01 09:02` loses the second row entirely under
+`errors="coerce"`. That is real data loss that reports as a clean parse, which is why the row count
+and the `unparsed_timestamps.csv` dump above exist.
+
+Two more traps. `tz_convert` on a naive column raises; `tz_localize` first if the source had no
+offset, and if you do not know the source timezone, say so rather than assuming UTC. And a
+`resample` over a sparse series fabricates the zeros it displays, which is correct for volume and
+wrong for anything you then average.
+
+### networkx
+
+Build and measure the graph in code, then hand it to Gephi to look at. Doing the metrics here
+rather than in the GUI is what makes them reproducible: the numbers come out of a script you can
+re-run and diff, instead of a sequence of panel clicks nobody recorded.
```bash
-pipx install datasette sqlite-utils
+pipx install networkx
+```
-sqlite-utils insert data.db records collected.csv --csv
-datasette data.db
-datasette publish vercel data.db --project my-investigation
+```python
+import networkx as nx
+import pandas as pd
+
+edges = pd.read_csv("edges.csv") # source, target, weight, kind
+G = nx.from_pandas_edgelist(edges, "source", "target",
+ edge_attr=["weight", "kind"], create_using=nx.DiGraph)
+print(G.number_of_nodes(), "nodes,", G.number_of_edges(), "edges")
+
+# who is most connected, in and out separately — they mean different things
+print(sorted(G.in_degree(weight="weight"), key=lambda x: -x[1])[:10])
+print(sorted(G.out_degree(weight="weight"), key=lambda x: -x[1])[:10])
+
+# who bridges otherwise separate clusters: usually the actual finding
+btw = nx.betweenness_centrality(G, normalized=True)
+print(sorted(btw.items(), key=lambda x: -x[1])[:10])
+
+# communities, computed on the undirected view as modularity requires
+communities = nx.community.greedy_modularity_communities(G.to_undirected())
+print([len(c) for c in communities])
+
+# is the graph one thing or several disconnected pieces
+print([len(c) for c in nx.weakly_connected_components(G)])
+
+# the ownership or reply chain between two specific entities
+print(nx.shortest_path(G, "acct_a", "acct_z"))
+
+# a subgraph two hops out from one node, which is what you actually draw
+ego = nx.ego_graph(G, "acct_a", radius=2)
+print(ego.number_of_nodes(), "nodes in the two-hop neighbourhood")
+
+# write metrics back onto the nodes, then export for Gephi
+nx.set_node_attributes(G, btw, "betweenness")
+nx.set_node_attributes(G, dict(G.degree()), "degree")
+nx.write_gexf(G, "case.gexf")
```
-`sqlite-utils` alone is worth learning — it handles the CSV-to-database step that otherwise eats an
-afternoon.
+Compute the metrics on the graph you can defend, not the one you collected. Betweenness on a graph
+built from "both accounts used the same hashtag" measures hashtag popularity and nothing else;
+betweenness on a graph of declared company directorships measures something real. The edge
+definition is the entire analysis, and it belongs in the caption.
+
+`betweenness_centrality` is roughly O(nodes × edges) and becomes impractical in the tens of
+thousands of nodes — pass `k=500` to sample pivots and accept an approximation. Note that
+`greedy_modularity_communities` needs the undirected view, and that community detection is
+stochastic in most implementations: run it twice, and if the communities differ materially, the
+structure is not there.
+
+### Gephi
+
+The place to look at the graph, after the metrics are computed. Current line is 0.11, free and
+open source, and it reads the `.gexf` written above with node attributes intact.
+
+```text
+1. File → Open → case.gexf. Check the import report for discarded edges.
+2. Data Laboratory tab first, not Overview. Confirm the node and edge counts
+ match what networkx printed. A mismatch means parallel edges were merged.
+3. Appearance → Nodes → Size → Ranking → betweenness (the attribute you wrote
+ from networkx). Size by the metric; never size by eye.
+4. Appearance → Nodes → Colour → Partition → modularity_class if you computed
+ it in Gephi, or your own community column from networkx.
+5. Layout → ForceAtlas2. Enable "Prevent Overlap" and "LinLog mode" only once
+ it has settled. Stop it; do not let it run to a shape you like.
+6. Filters → Topology → Degree Range to drop the degree-1 fringe, which is
+ usually most of the nodes and none of the finding.
+7. Preview tab for the export. Turn node labels on only for the nodes you will
+ name in the text.
+```
-## Timelines and maps
+The honest use of Gephi is to communicate a structure you already established numerically. A
+force-directed layout is not a measurement: the distance between two nodes on screen has no units,
+clusters appear in random data, and re-running the layout moves everything. If the figure is doing
+the arguing, the figure is overclaiming — put the degree and betweenness numbers in the caption and
+let the picture be an illustration.
+
+### Datawrapper
+
+Publishing a chart that survives being read on a phone by someone who does not trust you. It
+handles the things hand-rolled charts get wrong — colour-blind-safe palettes, responsive layout,
+accessible axis labelling, a visible source line — and the output is an embed plus a static
+fallback. Free tier covers everything below; the paid tiers add custom themes and private
+workspaces.
+
+```text
+1. Upload: paste the derived CSV (timeline.csv, not collected.csv) into the
+ Upload Data step. Check the "Check & Describe" screen — it shows the parsed
+ type per column, and this is where a month-first date misread surfaces
+ before it reaches the chart.
+2. Visualize → chart type. Lines for a rate over time, bars for counts by
+ category, a symbol map only if position is the finding.
+3. Refine → Appearance: leave the default palette. Its colours are chosen for
+ colour-vision deficiency and for greyscale printing.
+4. Refine → Axes: a bar chart's value axis starts at zero. Non-zero baselines
+ belong only on line charts, and then with the baseline labelled.
+5. Annotate → Title, description, source and byline. The source field is not
+ optional: name the dataset, the date collected, and the number of records.
+6. Annotate → highlight the specific data points you discuss in the text, and
+ add a shaded range for every known collection gap.
+7. Publish & Embed → take both the responsive iframe and the static PNG. The
+ PNG is what survives in a PDF, an email and an archive of your own article.
+```
-- **[Time.Graphics](https://time.graphics/)** — quick shareable timelines.
-- **[Pinpoint](https://journaliststudio.google.com/pinpoint/)** — OCR and entity extraction across
- large document sets; finds names and places across thousands of scanned pages.
-- **QGIS** — for anything where position matters. See
- [Maps & Satellite Imagery](/sheets/osint/maps-and-satellite-imagery).
+The failure mode is not the tool, it is reaching for a chart too early. A figure built from four
+observations is a figure about four observations, and "50% increase" on a base of four is noise
+rendered at 300 dpi. State the denominator in the subtitle, every time.
-## Publishing figures
+### Maps and spatial data
-**Datawrapper** produces clean, accessible, responsive charts and maps with almost no effort, and
-handles the things hand-rolled charts get wrong — colour-blind-safe palettes, mobile layout, proper
-axis labelling. **RAWGraphs** covers the less common chart types.
+Anything where position is the finding belongs in QGIS and GDAL, which are covered properly on
+[Maps & Satellite Imagery](/sheets/osint/maps-and-satellite-imagery) — including `qgis_process`
+for scripted geoprocessing and the `gdal_translate` / `gdalwarp` workflow. Use that sheet rather
+than a second copy here. The one rule that belongs on this page: a coordinate column is not spatial
+data until you know its CRS, and measuring distances in degrees because the layer is in `EPSG:4326`
+is the standard way to publish a wrong number.
-Whatever you use: label the axes, state the source, show the sample size, and do not start a bar
-chart's axis anywhere but zero.
+For document sets rather than coordinates,
+[Pinpoint](https://journaliststudio.google.com/pinpoint/) does OCR, transcription and entity
+extraction across thousands of scanned pages, and is free for journalists and researchers on
+application. It is the fastest route from a PDF dump to a searchable corpus, and its entity
+extraction is a lead generator, not a finding.
## Tool reference
@@ -129,6 +519,139 @@ chart's axis anywhere but zero.
- **Pretty visualisations oversell weak data.** The more convincing the figure, the more carefully
the caveats need stating.
- **Small numbers do not support percentages.** "50% increase" on a base of four is noise.
+- **Type inference is a silent rewrite.** csvkit reads `01/03/2026` month-first and emits
+ `2026-01-03`; pandas without `format="mixed"` coerces every value that does not match the first
+ row's format to `NaT`. Load as text, convert deliberately, and print the row count either side.
+- **A null is not a zero.** Columns with missing values make `COUNT(*)` and `COUNT(col)` disagree,
+ and a mean computed over the non-nulls reported against the full row count understates by exactly
+ the share that was missing.
+
+## Worked example
+
+One datum: `collected.csv`, 12,400 rows scraped from a set of accounts, with columns
+`id,ts,actor,company,amount,url`. The numbers below are invented; the order of the steps, and what
+each one reveals or destroys, is not.
+
+```bash
+# 1. before anything: does it parse, and how many rows are there really
+wc -l collected.csv
+csvcut -n collected.csv
+csvclean -a --label - collected.csv > /dev/null
+# 12401 collected.csv
+# 1: id 2: ts 3: actor 4: company 5: amount 6: url
+# 14 rows were longer than the header row
+```
+
+Fourteen ragged rows. Those are almost always unescaped commas inside a free-text field, which
+means the data in them is shifted one column right — `amount` holding a URL fragment. Fix or
+exclude them now, because every count below would otherwise be wrong by fourteen in an unknown
+direction.
+
+```bash
+# 2. the profile, with inference OFF so nothing is reinterpreted on the way in
+csvstat --nulls collected.csv
+csvstat -c company --freq --freq-count 15 collected.csv
+# amount: True (1,902 nulls)
+# { "Northgate Ltd": 412, "Northgate Limited": 198, "northgate ltd": 61,
+# "NORTHGATE LTD.": 44, "Westvale PLC": 390, ... }
+```
+
+Four spellings of one company, 715 rows between them. Counted raw, Northgate is the fourth-largest
+counterparty; counted merged, it is the largest. That single merge changes the headline, which is
+exactly why clustering comes before counting.
+
+```text
+# 3. OpenRefine: cluster the company column
+Edit cells → Cluster and edit → key collision / fingerprint
+ → 1 cluster, 4 values, 715 rows → merge to "Northgate Ltd"
+Then: key collision / ngram-fingerprint (n=2)
+ → 1 further cluster: "Westvale PLC" + "Westvale P.L.C" (390 + 7 rows)
+Then: nearest neighbour / levenshtein, radius 2
+ → proposes "Northgate Ltd" + "Northgate Mining Ltd" — REJECT.
+ Different companies. Export the cleaning operations as JSON so the
+ merge is reproducible and the rejection is on the record.
+```
+
+The levenshtein rejection is the part worth recording. An automatic merge at that radius would have
+folded two real companies into one and inflated the headline figure by another 130 rows, and nothing
+downstream would have shown it.
+
+```bash
+# 4. into SQLite, as text first, so the ambiguous dates survive the trip
+sqlite-utils insert case.db payments cleaned.csv --csv --pk id --no-detect-types
+sqlite-utils schema case.db
+sqlite-utils case.db "SELECT company, COUNT(*) n, COUNT(amount) with_amount
+ FROM payments GROUP BY 1 ORDER BY n DESC LIMIT 5"
+# [{"company": "Northgate Ltd", "n": 715, "with_amount": 601},
+# {"company": "Westvale PLC", "n": 397, "with_amount": 397}, ...]
+```
+
+114 Northgate rows have no amount. If you had summed without checking, the mean would have been
+computed over 601 rows and reported over 715 — a 19% understatement presented as a fact.
+
+```python
+# 5. the timeline, with the parse failures written out rather than dropped
+import pandas as pd
+df = pd.read_csv("cleaned.csv", dtype=str)
+df["ts"] = pd.to_datetime(df["ts"], errors="coerce", utc=True, format="mixed")
+print(df["ts"].isna().sum(), "of", len(df), "unparseable") # 38 of 12386
+df.loc[df["ts"].isna(), ["id", "ts", "url"]].to_csv("unparsed_timestamps.csv", index=False)
+df = df.dropna(subset=["ts"]).sort_values("ts")
+print(df.set_index("ts").resample("1D").size().describe())
+# the 38 all carry a trailing " (edited)" in the ts field — a scraper bug, recoverable
+```
+
+Thirty-eight rows lost to a scraper artefact, now visible in a file with their URLs, so they can be
+re-collected rather than quietly vanishing. Had `format="mixed"` been omitted, the count would have
+been in the thousands and would have looked equally clean.
+
+```python
+# 6. the actual claim: who posts within a minute of whom
+cols = ["ts", "actor", "id"]
+pairs = pd.merge_asof(
+ df.loc[:, cols].rename(columns={"actor": "a", "id": "id_a"}),
+ df.loc[:, cols].rename(columns={"actor": "b", "id": "id_b"}),
+ on="ts", tolerance=pd.Timedelta("60s"), direction="nearest", allow_exact_matches=False,
+)
+print(pairs.query("a != b").groupby(["a", "b"]).size().sort_values(ascending=False).head(10))
+# acct_f acct_k 94
+# acct_k acct_f 91
+# acct_f acct_m 12
+```
+
+`acct_f` and `acct_k` land within sixty seconds of each other 94 times. That is the finding, and it
+is a sentence with a number in it. The graph comes next only to show *where* those two sit, not to
+establish that they are linked.
+
+```python
+# 7. the graph, with metrics computed before anything is drawn
+import networkx as nx
+edges = pairs.query("a != b").groupby(["a", "b"]).size().reset_index(name="weight")
+G = nx.from_pandas_edgelist(edges, "a", "b", edge_attr="weight", create_using=nx.DiGraph)
+btw = nx.betweenness_centrality(G, normalized=True)
+print(sorted(btw.items(), key=lambda x: -x[1])[:5])
+nx.set_node_attributes(G, btw, "betweenness")
+nx.write_gexf(G, "case.gexf")
+# acct_f 0.31, acct_k 0.28, acct_m 0.04, ...
+```
+
+```bash
+# 8. publish the queryable dataset and the figure
+datasette serve -i case.db -m metadata.yml -o
+```
+
+Then `case.gexf` into Gephi, sized by the `betweenness` attribute already on the nodes, degree-1
+fringe filtered out, labels on for `acct_f` and `acct_k` only. And the daily volume series into
+Datawrapper, with the two-day collection gap on 5–7 March drawn as a shaded band and the source
+line reading "12,386 posts collected 2026-02-01 to 2026-03-31; 38 timestamps unparsed; four company
+name variants merged".
+
+What this run establishes: a corrected counterparty ranking, a reproducible cleaning record
+including one rejected merge, and a quantified timing relationship between two accounts. What it
+does not establish: that `acct_f` and `acct_k` are operated by the same person, or coordinated at
+all — posting within a minute is consistent with coordination, with both reacting to the same
+trigger, and with a scheduler neither of them controls. The chart cannot distinguish those, and
+saying so is the finding's credibility.
## Broader catalogues
diff --git a/src/content/sheets/osint/geolocation.md b/src/content/sheets/osint/geolocation.md
@@ -4,9 +4,9 @@ description: "Work out where a photo was taken, and when, from shadows, sun posi
category: osint
subcategory: "Geospatial"
tags: [osint, geolocation, chronolocation, shadows, verification]
-tools: [suncalc, shadowmap, geohints, qgis]
+tools: [suncalc, shadowfinder, pysolar, shadowmap, shademap, geohints, qgis, gdal, exiftool]
difficulty: advanced
-updated: 2026-09-28
+updated: 2026-10-04
references:
- name: "Bellingcat's Online Investigation Toolkit"
url: "https://bellingcat.gitbook.io/toolkit"
@@ -24,9 +24,10 @@ references:
## What this covers
-Placing an image on the earth and in time without metadata. This is the discipline OSINT is best
-known for and the most labour-intensive thing in it: a hard geolocation is hours of work, not
-minutes.
+Placing an image on the earth and in time without metadata. The place comes from matching what is
+in the frame against the world; the time comes from the sun, which is the one thing in the picture
+whose behaviour is exactly calculable. A hard geolocation is hours of work, and the chronolocation
+that follows it is arithmetic you can show someone.
## Method
@@ -36,50 +37,367 @@ minutes.
2. **Narrow the region.** Those details constrain the country or region long before they give you a
point. Pole and bollard design alone often gets you to a handful of countries.
3. **Find a searchable anchor.** A business name, a phone number on a van, a street name, a bus
- route number. One readable sign collapses the search.
-4. **Match terrain.** Mountain ridgelines are effectively fingerprints and are visible from far
- away. Compare against a terrain viewer.
-5. **Confirm with imagery.** Satellite and street view, ideally from the era of the photo. Check
- building footprints, not just the general scene.
-6. **Chronolocate.** Shadow direction and length plus a known position gives you a time of day and
- a range of dates.
+ route number. One readable sign collapses the search; spend your effort on making one legible
+ before you spend it on scanning imagery.
+4. **Match terrain.** Ridgelines are fingerprints and are visible from tens of kilometres away.
+ This is the step that works when there is no signage at all, and the step that fails on flat
+ ground.
+5. **Confirm with imagery.** Satellite and street view from the era of the photo. Match building
+ footprints and the relative geometry of fixed objects, not the general vibe of the scene.
+6. **Measure the shadow, then compute.** Direction gives you the sun's azimuth, the length-to-height
+ ratio gives you its elevation. Together they give a time of day and, usually, two candidate date
+ windows per year rather than one.
+7. **Say what would falsify it.** A geolocation you cannot break is one nobody checked. Name the
+ feature that would have to be in the wrong place for your answer to be wrong.
-## Shadow and sun work
+The judgement calls: decide early whether you are doing signage work or terrain work, because they
+need different tools and mixing them wastes hours. Decide whether the date is given or derived —
+if you take the date from the caption and then use the sun to "confirm" the caption, you have
+proved nothing. And decide whether a precise coordinate should be published at all, before you
+find one.
-Given a location and a date, the sun's position is exactly calculable — so a shadow in a photo
-constrains the time it was taken, and if you know the time it constrains the date.
+## The shadow arithmetic
-| Tool | What it does |
-| --- | --- |
-| [SunCalc](https://suncalc.org/) | Sun position, shadow direction and length for any place, date and time. The workhorse. |
-| [ShadowMap](https://shadowmap.org/) | 3D buildings with cast shadows rendered at a chosen time. |
-| [ShadeMap](https://shademap.app/) | Global shadow simulation including terrain and trees. |
-| [Shadow Finder](https://github.com/bellingcat/ShadowFinder) | Inverts the problem: given shadow length and a timestamp, maps every point on earth where that shadow is possible. |
+A vertical object of height `h` casting a shadow of length `s` on level ground fixes the sun's
+elevation angle:
+
+```text
+tan(elevation) = h / s
+
+h = 5.00 m, s = 8.20 m
+elevation = atan(5.00 / 8.20) = atan(0.60976) = 31.37 degrees
+
+check: 5.00 / tan(31.37 deg) = 5.00 / 0.60976 = 8.20 m
+```
+
+The shadow points directly away from the sun, so the sun's compass bearing is the shadow's bearing
+plus 180:
+
+```text
+shadow bearing measured off the image : 062 deg
+sun azimuth : (062 + 180) mod 360 = 242 deg
+```
+
+Three things break this. The ground must be level — a shadow running downhill is longer than the
+arithmetic expects and will push your elevation too low. The object must be vertical and you must
+be measuring its true height, not its foreshortened height in the frame. And `elevation` here is
+geometric; at elevations below about 5 degrees atmospheric refraction lifts the apparent sun by
+roughly half a degree, which matters at sunrise and sunset and nowhere else.
+
+Then the part people get wrong: an azimuth and elevation pair does **not** identify one date. The
+sun retraces its declination on the way out of winter and on the way back in, so almost every
+pair matches two windows in the year, roughly symmetric about the solstice. Expect two answers and
+rule one out with something other than the sun — foliage, snow, an event in the background, the
+clothes people are wearing.
+
+## Key tools
+
+### SunCalc
+
+Web only, free, no account. Give it a coordinate and a date and it draws the sun's track for that
+day with the numbers underneath: altitude, azimuth, and a **shadow length** readout for an object
+height you type in. That last field is why it beats the generic ephemeris sites — you can go
+straight from the measurement you took off the image to the time that produces it, without
+converting anything.
+
+```text
+1. Enter the coordinate in "Set Lat/Lon", or drag the map pin onto the spot.
+2. Set month and date in the dropdowns. Note the timezone line (TZ) it reports; everything
+ downstream is in that zone, and it is the usual source of a one-hour error.
+3. Drag the time slider and watch two readouts: "Azimuth" against your measured sun bearing,
+ and "Shadow length [m]" against the shadow you measured, with "at an object level [m]"
+ set to the object's real height.
+4. When both match at once, you have a time. When only the azimuth matches, you have two
+ times -- morning and afternoon mirror each other, and length is what separates them.
+5. "Reverse Calculation" works the other way: fix the shadow and ask for the time.
+6. Capture the page URL. The state lives in the fragment --
+ #/<lat>,<lon>,<zoom>/<YYYY.MM.DD>/<HH:MM>/<object height> -- so the link reproduces the
+ exact reading rather than just the tool.
+```
+
+The azimuth it reports is a compass bearing measured clockwise from north, which is what you
+measured off the image, so no conversion is needed here. Read-offs from a slider are worth about a
+minute of precision at best; quote a window, not a timestamp. And the whole result is conditional
+on the coordinate — a 50 m position error barely moves the azimuth, but a wrong *date* moves it by
+degrees per week away from the solstices, which is exactly where the two-candidate-window problem
+bites.
+
+### ShadeMap and Shadowmap
+
+Also web only. These solve the opposite problem: when the shadow in your image is cast by a
+building or a ridge rather than by the subject, you cannot measure an object height, so SunCalc's
+arithmetic has nothing to chew on. These render the shadows that real geometry casts at a chosen
+moment and you match the *pattern* instead.
+
+```text
+ShadeMap (shademap.app)
+ - loads OpenStreetMap building footprints plus a terrain model, so both urban and
+ mountain shadows are simulated
+ - date and time slider at the bottom; the shadow polygons redraw live
+ - the sun-exposure analytics (hourly / daily / annual) are a solar-study feature,
+ not an investigative one -- ignore them for this work
+ - free for interactive use on the site; the paid tiers are for embedding the engine
+ in your own map, not for using it here
+
+Shadowmap (shadowmap.org)
+ - 3D buildings with cast shadows, better for dense city blocks
+ - set the date and time, then orbit the camera to the photographer's approximate
+ position and compare the shadow edges against the frame
+ - sign-in gates some of the time controls
+
+What to capture either way: a screenshot at the matched time, the coordinate, the
+date, the stated timezone, and the building or ridge you matched on.
+```
+
+Building shadows are only as good as the building heights in OpenStreetMap, which are frequently
+absent, guessed, or recorded for a structure that has since been extended. Terrain shadows are more
+trustworthy because the DEM is measured. Neither knows about trees, awnings, scaffolding or a
+crane that was on site that month — so an unexplained shadow is a reason to doubt your position
+before it is a reason to doubt the tool.
+
+### suncalc (Python)
+
+The same model as the website, as a library, which is what you want once you are checking more than
+one candidate. Give it a coordinate and a timestamp and it returns the sun's position; loop it over
+a year and you get every moment that matches your measurement, which is the honest way to find both
+date windows instead of stopping at the first one.
+
+```bash
+pipx install suncalc # or: pip install suncalc pandas
+```
+
+```python
+import math, datetime
+import pandas as pd
+from suncalc import get_position, get_times
+
+lat, lon = 38.7075, -9.1364 # candidate coordinate
+height, shadow = 5.00, 8.20 # metres, measured off the image
+shadow_bearing = 62.0 # degrees, measured off the image
+
+# the two numbers the image gives you
+elevation = math.degrees(math.atan(height / shadow)) # 31.37
+sun_bearing = (shadow_bearing + 180) % 360 # 242.0
+
+def sun(dt):
+ # suncalc returns radians, and its azimuth is measured from SOUTH toward west --
+ # add 180 degrees to get the compass bearing you measured off the image
+ p = get_position(pd.Timestamp(dt), lon, lat)
+ return (math.degrees(p["azimuth"]) + 180) % 360, math.degrees(p["altitude"])
+
+# sweep the whole year at 5-minute resolution: this is what exposes the second window
+start = datetime.datetime(2026, 1, 1, tzinfo=datetime.timezone.utc)
+for step in range(0, 365 * 288):
+ dt = start + datetime.timedelta(minutes=5 * step)
+ az, alt = sun(dt)
+ if abs(az - sun_bearing) < 1.0 and abs(alt - elevation) < 0.5:
+ print(f"{dt:%Y-%m-%d %H:%M} UTC az {az:6.1f} elev {alt:5.1f}")
+
+# sanity-check the day itself: if your candidate time is after dusk, the coordinate is wrong
+t = get_times(pd.Timestamp("2026-03-22 12:00:00"), lon, lat)
+print("sunrise", t["sunrise"], "solar noon", t["solar_noon"], "sunset", t["sunset"])
+```
+
+Two conventions to hold onto, because getting either wrong silently produces a plausible, wrong
+answer. Azimuth comes back in radians measured from south, so the `+ 180` above is not optional.
+Altitude is geometric and ignores refraction. Pass timezone-aware UTC datetimes and convert to
+local time at the very end, after you have checked whether the location observes DST on that date
+— a shadow matched in March against a September local time is the most common way this analysis
+goes quietly wrong. The package is thin and has not changed since 2023; that is fine for a model
+of celestial mechanics, and it is not a sign of anything.
+
+### pysolar
+
+A second implementation, which is the point of using it: run the same coordinate and timestamp
+through both and you are checking your arithmetic rather than one library's. It is also the more
+natural fit when you want irradiance or want to sweep without pandas in the way.
+
+```bash
+pipx install pysolar # or: pip install pysolar
+```
+
+```python
+import datetime
+from pysolar.solar import get_altitude, get_azimuth
+
+lat, lon = 38.7075, -9.1364
+when = datetime.datetime(2026, 6, 15, 12, 30, tzinfo=datetime.timezone.utc)
+
+# pysolar's azimuth is already a compass bearing, clockwise from north -- no offset needed
+print("altitude", get_altitude(lat, lon, when)) # 56.80
+print("azimuth ", get_azimuth(lat, lon, when)) # 216.45
+
+# cross-check against suncalc for the same instant: 56.79 / 216.36
+# agreement to a tenth of a degree means the inputs are right; a whole-degree gap
+# almost always means one of the two got a naive datetime and assumed local time
+```
+
+`get_altitude` applies a refraction correction by default, which is why it sits a hundredth of a
+degree off suncalc's geometric value near the zenith and further off near the horizon — at 5
+degrees elevation the two will differ by about half a degree, and pysolar is the one closer to what
+a camera actually saw. Passing a naive `datetime` is the failure mode: it will be interpreted as
+UTC and you will get an answer that looks fine and is hours wrong.
-Shadow Finder is the one worth knowing about, because it works when you have *no* candidate
-location:
+### ShadowFinder
+
+Bellingcat's inversion of the problem, for when you have no candidate coordinate at all. Feed it an
+object height, a shadow length and a UTC timestamp and it maps every point on earth where a shadow
+of that ratio is possible at that instant — a band, not a point, but a band you can intersect with
+everything else you know.
+
+```bash
+pipx install shadowfinder # or: pip install shadowfinder
+
+# object height, shadow length, date, time -- POSITIONAL, in that order
+shadowfinder find 5 8.2 2026-03-22 16:00:00 --time_format=utc
+
+# when you already have the sun's elevation rather than a height/length pair
+shadowfinder find_sun 31.37 2026-03-22 16:00:00 --time_format=utc
+
+# the subcommands and their arguments, before you guess
+shadowfinder find --help
+shadowfinder find_sun --help
+
+# writes a PNG map into the current directory, named for its inputs:
+# shadow_finder_20260322-160000-Utc_5_8.2.png
+ls shadow_finder_*.png
+
+# the timezone grid it uses for local-time input is generated once and cached
+shadowfinder generate_timezone_grid
+```
+
+The older `--object-height / --shadow-length / --date-time` flag form that circulates in tutorials
+no longer exists — current releases take positional arguments under a `find` subcommand and will
+print a bare usage line if you use the old syntax. The output band is wide, and it is a band of
+*possibility*: it tells you where such a shadow could fall, not where this one did. It is at its
+most useful when the timestamp is independently known — from a livestream, a broadcast, a
+timestamped post — and least useful when you derived the timestamp from the same photo.
+
+### GeoHints
+
+Web only, free, and the reference for step one. It catalogues the regional giveaways by country:
+bollards, utility poles, traffic lights, road markings, licence plates, house numbers, sign shapes,
+road-numbering schemes, post boxes, driving side, even the vehicles Google mounts its cameras on.
+Built for GeoGuessr, which is precisely why it is organised the way an investigator needs — by
+visible object rather than by country.
+
+```text
+1. Pick the object in your frame that is most likely to be regulated nationally.
+ Ranked by how much they narrow things: licence plate format > road sign shape and
+ font > utility pole construction > bollard > kerb and road marking > house numbers.
+2. Open that object's page and scan the per-country plates side by side. You are
+ looking for a disqualifier as much as a match -- "not this country" eliminates
+ faster than "maybe this one".
+3. Cross two independent objects before you commit to a region. A pole design shared
+ by six countries and a sign font shared by four may intersect at one.
+4. Check the Signs subsection separately: back-of-sign colour, chevron style and
+ bus-stop design are all national and all usually visible in a frame that has no
+ readable text at all.
+5. Record which object you matched on and the page you matched it against. "Bollard
+ type matches Portugal" is checkable; "looks Iberian" is not.
+```
+
+It is a crowd-built reference, so coverage is uneven — rich for Europe and the GeoGuessr-popular
+countries, thin for much of Africa and Central Asia, and it lags on recent sign redesigns and plate
+reissues. Nothing on it is evidence by itself: a match narrows the search space, and the actual
+geolocation still has to end at a specific place on a specific image.
+
+### QGIS and gdal_viewshed
+
+Terrain is the fallback when there is no text in the frame. A ridgeline's silhouette depends only
+on the observer's position, so if you can trace the horizon in the photo you can test candidate
+positions against a DEM until one produces that profile. QGIS has carried a native **Elevation
+Profile** panel since 3.26, which is the fastest way to do it interactively.
+
+```text
+QGIS elevation profile, for matching a horizon
+1. Load a DEM (SRTM, Copernicus GLO-30, or a national LiDAR product where one exists).
+2. Layer Properties > Elevation > tick "Represent elevation surface". Without this the
+ profile panel will not see the layer and the panel looks broken.
+3. View > Elevation Profile to open the panel.
+4. "Capture curve": draw a line from your candidate camera position out along the
+ bearing the photograph faces, past the furthest ridge.
+5. Compare the plotted skyline against the photo's horizon. Use the vertical
+ exaggeration control to match the photo's apparent relief, then set it back to 1
+ before you quote any number off it.
+6. Wrong candidate positions fail obviously -- a peak in the wrong order, a saddle
+ that should be hidden. That is the point: this step eliminates fast.
+```
+
+For the reverse question — "could this spot have been seen from there at all" — `gdal_viewshed`
+answers it as a raster, which is better than eyeballing when you have many candidates:
+
+```bash
+# binary visibility from an observer 2 m above the surface, out to 15 km
+# -ox/-oy are in the DEM's own CRS units, so reproject the DEM to a metric CRS first
+gdal_viewshed -md 15000 -ox 499000 -oy 4283000 -oz 2 dem_utm.tif viewshed.tif
+
+# the minimum height a target would need to be visible, rather than a yes/no mask
+gdal_viewshed -om DEM -md 15000 -ox 499000 -oy 4283000 dem_utm.tif minheight.tif
+
+# accumulate over a grid of observers: where could a photographer have stood
+gdal_viewshed -om ACCUM -md 15000 -ox 499000 -oy 4283000 dem_utm.tif accum.tif
+
+# target height matters when the subject is a mast or a roofline, not the ground
+gdal_viewshed -md 20000 -tz 30 -ox 499000 -oy 4283000 dem_utm.tif mast_visible.tif
+```
+
+A DEM is a bare-earth or surface model depending on which one you downloaded, and the difference
+decides the answer: a bare-earth model will happily tell you that you can see through a forest and
+a city. 30 m postings smooth real ridgelines, so a profile that is close but not exact is expected
+and is not a match. And viewshed output is geometric — it knows nothing about haze, which is what
+actually limits how far a camera sees on the day in question. Reprojection, cropping and
+hillshading of these rasters are covered on
+[Maps, Satellite & Street-Level Imagery](/sheets/osint/maps-and-satellite-imagery).
+
+### exiftool
+
+Before any of the above, check whether the question is already answered. Most images stripped by
+social platforms have nothing, but files received directly from a source, pulled from a cloud
+drive, or lifted off a messaging app's media folder frequently still carry GPS.
```bash
-pip install shadowfinder
-shadowfinder --object-height 2 --shadow-length 3.5 \
- --date-time "2024-06-15 14:30:00"
+brew install exiftool # or: apt install libimage-exiftool-perl
+
+# GPS as a decimal pair you can paste into a map, and nothing else
+exiftool -n -GPSLatitude -GPSLongitude -GPSAltitude photo.jpg
+
+# the direction the camera was pointing, which turns a point into a line of sight
+exiftool -GPSImgDirection -GPSImgDirectionRef photo.jpg
+
+# every timestamp in the file, including the offset that tells you the local zone
+exiftool -time:all -OffsetTime -OffsetTimeOriginal -a -G1 photo.jpg
+
+# a whole directory into one CSV, to find which file in a set still has coordinates
+exiftool -csv -n -GPSLatitude -GPSLongitude -DateTimeOriginal ./images > gps.csv
```
-You need the object's height and its shadow's length in the same units, and a timestamp. It returns
-a band of possible latitudes.
+A GPS tag is a claim made by the device, not a fact: it can be a cached fix from the last place the
+phone had signal, it can be the location the file was edited rather than shot, and it can be
+written by hand. Treat it as a strong lead that still has to be confirmed against the frame. Full
+metadata, container and manipulation analysis lives on
+[Image & Video Forensics](/sheets/osint/image-video-forensics) rather than here.
+
+## Shadow and sun tools at a glance
+
+| Tool | What it does |
+| --- | --- |
+| [SunCalc](https://www.suncalc.org/) | Sun position, azimuth, altitude and shadow length for any place, date and time. The workhorse. |
+| [ShadowMap](https://shadowmap.org/) | 3D buildings with cast shadows rendered at a chosen time. |
+| [ShadeMap](https://shademap.app/) | Global shadow simulation including terrain as well as buildings. |
+| [Shadow Finder](https://github.com/bellingcat/ShadowFinder) | Inverts the problem: given shadow length and a timestamp, maps every point on earth where that shadow is possible. |
+| [suncalc-py](https://github.com/kylebarron/suncalc-py) | The SunCalc model as a Python library, for sweeping many candidates. |
+| [pysolar](https://pysolar.readthedocs.io/) | Independent solar-position implementation, for cross-checking. |
## Visual reference
[GeoHints](https://geohints.com/) catalogues the regional details that narrow a location — bollards,
-traffic lights, road markings, utility poles, licence plates — by country. It was built for
-GeoGuessr and is genuinely the best reference for this step.
+traffic lights, road markings, utility poles, licence plates — by country.
[Bellingcat's OpenStreetMap Search](https://osm-search.bellingcat.com/) and
-[Spot](https://spot.bellingcat.com/) both let you search for *relationships* between features —
+[Spot](https://www.findthatspot.io/) both let you search for *relationships* between features —
"a church within 200m of a bridge over a river" — which is how you turn a described scene into
-candidate coordinates.
-
-Further mapping, satellite and street-level tools are listed on
+candidate coordinates. The Overpass queries underneath that idea are on
[Maps, Satellite & Street-Level Imagery](/sheets/osint/maps-and-satellite-imagery).
## Tool reference
@@ -96,13 +414,74 @@ Further mapping, satellite and street-level tools are listed on
- **Satellite imagery has a date.** A building present in the photo and absent from imagery may
simply be newer, or demolished. Check the capture date and look for historical imagery.
-- **Terrain matching defeats you at low elevation.** Flat terrain has no ridgeline to match.
-- **Shadow work needs the true timezone**, including DST, and the analysis is only as good as your
- measurement of the shadow.
+- **Terrain matching defeats you at low elevation.** Flat terrain has no ridgeline to match, and a
+ 30 m DEM will not reproduce a 10 m rise.
+- **Shadow work needs the true timezone**, including whether DST was in force on that date in that
+ country, which is a different question from whether it is in force now.
+- **One azimuth-and-elevation pair means two date windows.** Reporting the first one you find, and
+ not the second, is the most common way a chronolocation is wrong while being arithmetically
+ correct.
+- **Circular reasoning.** If the date came from the caption, the sun cannot confirm the caption.
+ It can only confirm that the caption is internally consistent, which is a much weaker statement.
- **Publishing a precise location endangers people.** For conflict imagery especially, consider
whether the coordinates need to be public.
- **Plausible is not confirmed.** A location that fits is a hypothesis. Confirmation means a
- specific feature matching in a specific place.
+ specific feature matching in a specific place, named, so that someone else can go and disagree.
+
+## Worked example
+
+One datum: a photograph of a public square. No metadata, no caption beyond a claim that it was
+taken "in September". A lamp standard in the foreground casts a clean shadow across flat paving.
+
+1. **Inventory and narrow.** The bollards, the kerb profile and the black-on-white house numbers
+ go into GeoHints object by object. The sign font and the bollard type intersect on Portugal;
+ a partial bus-stop livery in the background does not contradict it.
+2. **Anchor.** A shopfront name is legible after a crop and upscale. Searching the name plus
+ "Lisboa" returns one address on a named square, and Street View from the same corner reproduces
+ the building line, the arcade spacing and the lamp standard. Candidate coordinate:
+ **38.7075, -9.1364**.
+3. **Measure.** The lamp standard matches a documented 5.00 m type and its shadow, scaled against
+ the paving slabs, runs 8.20 m. The shadow's bearing, taken off the satellite image along the
+ same paving joint, is 062 degrees.
+4. **Convert.** `atan(5.00 / 8.20)` = **31.37 degrees** solar elevation. Sun bearing =
+ `(62 + 180) mod 360` = **242 degrees**, so the sun is in the west-south-west and this is an
+ afternoon photograph.
+5. **Sweep the year.** The `suncalc` loop above, at the candidate coordinate, with a 1-degree
+ azimuth and 0.5-degree elevation tolerance, returns seven matching instants in all of 2026, in
+ two clusters:
+
+```text
+2026-03-21 16:00 UTC az 241.6 elev 31.0
+2026-03-22 16:00 UTC az 242.0 elev 31.2
+2026-03-23 16:00 UTC az 242.4 elev 31.5
+2026-03-24 16:00 UTC az 242.8 elev 31.7
+
+2026-09-20 15:45 UTC az 242.1 elev 31.8
+2026-09-21 15:45 UTC az 241.8 elev 31.5
+2026-09-22 15:45 UTC az 241.6 elev 31.1
+```
+
+6. **Convert to local time, carefully.** Portugal moved to UTC+1 on 29 March 2026 and back on 25
+ October. The March window is therefore **16:00 local**, and the September window is **16:45
+ local**. Same sun, different clock, entirely because of a date the arithmetic knows nothing
+ about.
+7. **Break the tie with something that is not the sun.** The trees on the square are in full leaf
+ and the café terrace is laid out — both argue September over the third week of March. The
+ caption's "September" is now consistent with the image rather than assumed by it, because the
+ sun was computed independently and the foliage was the tiebreaker.
+8. **Cross-check the model.** The same instant through `pysolar` returns an elevation within a
+ tenth of a degree of `suncalc`'s, which confirms the inputs rather than the conclusion.
+
+What you can assert: a square identified by a named shopfront and reproduced building geometry, and
+an afternoon sun position consistent with **20–22 September 2026, around 16:45 local time**, with
+21–24 March as a rejected alternative and the reason for rejecting it stated.
+
+What would falsify it: a shopfront that was at a different address on that date, paving that was
+relaid between the Street View capture and the photograph (which would break the shadow
+measurement, not the location), a lamp standard of a different height, or any evidence that the
+foliage argument is wrong. The measurement that carries the most risk is the 8.20 m — a 10 per cent
+error there moves the elevation by roughly 2.5 degrees and the window by several days, so quote it
+with that tolerance or do not quote a day at all.
## Broader catalogues
@@ -111,7 +490,7 @@ Further mapping, satellite and street-level tools are listed on
## More tools
-Further tools for this area from the OSINT Newsletter Tools Library (), excluding those already listed above.
+Further tools for this area from the OSINT Newsletter Tools Library ([Geolocation and Maps OSINT](https://tools.osintnewsletter.com/tool-categories/geolocation-and-maps-osint)), excluding those already listed above.
| Tool | What it does |
| --- | --- |
diff --git a/src/content/sheets/osint/image-video-forensics.md b/src/content/sheets/osint/image-video-forensics.md
@@ -196,8 +196,9 @@ ffmpeg -i video.mp4 -vf "select='gt(scene,0.3)',metadata=print:file=scenes.txt"
`-c copy` matters for anything you will publish or hand to someone else: re-encoding an excerpt
destroys the compression history that any later analysis would use, and makes your excerpt
-unfalsifiable in the wrong direction. `-fps_mode vfr` replaced `-vsync vfr` in ffmpeg 5; both are
-accepted on current builds but only the former is documented.
+unfalsifiable in the wrong direction. `-fps_mode vfr` replaced `-vsync vfr` in ffmpeg 5, and `-vsync`
+was removed outright in ffmpeg 9 — it now fails with `Unrecognized option 'vsync'`. Any older
+cheatsheet or Stack Overflow answer using it will not run.
Scene-change detection finds cuts, and cuts in material claimed to be a single continuous recording
are worth explaining. It also fires on pans, flashes and camera shake, so read the list as
diff --git a/src/content/sheets/osint/osint-foundations.md b/src/content/sheets/osint/osint-foundations.md
@@ -4,9 +4,9 @@ description: "How to run an open-source investigation without burning yourself o
category: osint
subcategory: "Foundations"
tags: [osint, methodology, opsec, verification]
-tools: [hunchly, obsidian, logseq]
+tools: [hunchly, obsidian, logseq, multipass, lima, docker, curl, wget, dig, shasum]
difficulty: beginner
-updated: 2026-09-28
+updated: 2026-10-04
references:
- name: "Bellingcat's Online Investigation Toolkit"
url: "https://bellingcat.gitbook.io/toolkit"
@@ -30,10 +30,11 @@ references:
## What this covers
-The habits that decide whether an investigation holds up: how you look at a target without
-telling them, how you record what you found so it survives the page being deleted, and how you
-avoid deciding the answer before you have it. Tooling is the easy part of OSINT. This is the part
-that separates a finding from a guess.
+The habits that decide whether an investigation holds up: how you look at a target without telling
+them, how you record what you found so it survives the page being deleted, and how you avoid
+deciding the answer before you have it. Most of this sheet is discipline rather than tooling,
+because most of it has no command — and where a tool genuinely exists, the commands below are the
+ones that keep you from leaking.
## The rule that matters most
@@ -44,19 +45,459 @@ that only three people have seen tells those three people someone is looking.
Decide before you start: is this target likely to notice, and does it matter if they do?
-## Collection hygiene
+## Method
-1. **Separate identity.** A research browser profile, or better a separate VM, with its own
- accounts. Never the account you use for anything else. Expect platforms to ban research
- accounts eventually and do not build anything you cannot lose.
-2. **Capture as you go, not afterwards.** Anything interesting gets archived the moment you see
- it. Pages disappear, get edited, or go private within hours of someone noticing attention.
-3. **Record the URL, the timestamp, and how you got there.** A screenshot with no source and no
- date is worth nothing. The path you took to a finding is part of the finding.
-4. **Keep raw separate from conclusions.** One place for what you collected, another for what you
- think it means. Conflating them is how an assumption becomes a fact three notes later.
-5. **Write down what you looked for and did not find.** Negative results stop you re-running the
- same dead end next week, and they are what an honest report needs.
+1. **Write the question down before you collect anything.** One sentence, answerable, with a
+ stated standard of proof. "Is this company linked to that one" is not a question; "does any
+ public filing show a shared officer, address or beneficial owner" is.
+2. **Decide your exposure budget first.** Which identity touches the target, from what address,
+ and what happens if that address is logged. Changing this mid-investigation is how an account
+ gets burned.
+3. **Capture as you go, never afterwards.** Anything interesting is archived the moment you see it.
+ Pages disappear, get edited or go private within hours of someone noticing attention. See
+ [Archiving & Evidence Preservation](/sheets/osint/archiving-and-evidence).
+4. **Record the route, not just the result.** URL, UTC timestamp, the query that surfaced it, and
+ what you clicked to get there. A screenshot with no source and no date is worth nothing, and the
+ path to a finding is the part you forget first.
+5. **Keep raw separate from conclusions, in different files.** One place for what you collected,
+ another for what you think it means. Conflating them is how an assumption becomes a fact three
+ notes later and a published claim three weeks later.
+6. **Log what you looked for and did not find.** Negative results stop you re-running the same
+ dead end next week, and they are what an honest report needs in order to describe its own limits.
+7. **Name the thing that would prove you wrong, then go look for it.** If you cannot state what
+ evidence would change your conclusion, you are not investigating, you are assembling.
+8. **Stop at the question you wrote down.** The collection will always offer you more. More about a
+ bystander, more about a family member, more that is interesting and none of your business.
+
+## Key tools
+
+### Hunchly
+
+The one tool that removes the discipline problem, because it captures continuously rather than when
+you remember to. It is a browser extension that records every page you visit while a case is open:
+the full HTML and resources, the URL, a UTC timestamp and a hash of the capture, into a local case
+file that is full-text searchable. That hash-at-capture-time is the evidentiary point — it is the
+same argument as the manifest pattern in
+[Archiving & Evidence Preservation](/sheets/osint/archiving-and-evidence), applied automatically to
+everything you looked at including the pages that turned out not to matter.
+
+Commercial, subscription, now owned by Maltego Technologies, with a 30-day trial that needs no card.
+There is no CLI and no API: this is a GUI tool and the workflow is the product.
+
+```text
+1. Install the extension, open the desktop app, and create a case BEFORE you
+ start looking. Captures outside an open case are not recorded.
+2. Set the case name to your case ID, not to the target's name. The case file
+ leaks the target's name to anyone who sees your screen or your backups.
+3. Toggle capture ON. The extension icon state is the only indication; check it
+ after every browser restart, because a silent-off session is unrecoverable.
+4. Add "selectors" — the names, handles, domains and phone numbers you are
+ tracking. Hunchly then highlights them on every page you visit and logs
+ which page each one appeared on. This is the feature people underuse: it
+ catches a name in a footer you would never have read.
+5. Tag pages as you go. An untagged case file of 4,000 pages is a search index,
+ not a narrative.
+6. Export: the case export carries the captured pages, the hashes, the timestamps
+ and the selector hit log. Export at milestones, not just at the end — the
+ case file is a single local database and a single local database can corrupt.
+7. Storage choice matters. Local keeps everything on your machine; the hosted
+ option puts your captures and therefore your whole research pattern on
+ someone else's infrastructure. Pick deliberately and write down which.
+```
+
+What it does not do: a Hunchly capture is a record that *you* saw that content at that time, with no
+third-party attestation. It is excellent provenance for your own process and weak evidence against
+a determined challenge, which is why anything load-bearing still gets a public-archive capture and
+an independent timestamp. It also captures everything, including pages you visited by accident and
+pages containing other people's personal data — treat the case file as sensitive material in its own
+right.
+
+### Obsidian or Logseq: the case vault
+
+Local, plain-text, file-per-entity notes with links between them. The reason this beats a document
+is that an investigation is a graph of entities, not a narrative, and the thing you need six weeks
+in is "every note that mentions this phone number". Obsidian suits entity-per-note work where the
+links between people are the finding; Logseq's daily journal and outliner suit chronology-heavy
+work. Both store Markdown on disk, so the vault is greppable, diffable and `git`-able without the
+application.
+
+Structure the vault so that provenance cannot be separated from content:
+
+```text
+case-2026-014/
+ 00-question.md the question, the standard of proof, the stop condition
+ 10-entities/
+ person-a-khan.md one file per entity, named by role not by guesswork
+ company-northgate.md
+ account-acct_f.md
+ 20-raw/ collected artefacts, never edited
+ 2026-10-04T0407Z-post123.warc.gz
+ 2026-10-04T0407Z-post123.info.json
+ MANIFEST.sha256
+ 30-findings/
+ f01-post-edited-after-upload.md
+ 40-negative/
+ not-found.md every search that returned nothing, with the query
+ 90-log.md append-only: what you did, when, from which identity
+```
+
+Every entity note carries its own provenance header, so a claim can never be read without its
+source:
+
+```markdown
+---
+entity: company-northgate
+type: company
+aliases: ["Northgate Ltd", "Northgate Limited", "northgate ltd"]
+confidence: medium
+first_seen: 2026-10-03
+last_checked: 2026-10-04
+---
+
+## Established
+
+- Registered number 09876543, UK. Source: Companies House API, retrieved
+ 2026-10-04T04:07Z. Raw: `20-raw/ch-09876543.json` (sha256 4b1c8e...).
+
+## Reported but unverified
+
+- Described as "the trading arm" in the filing at para 17. Single source,
+ uncorroborated, author has an interest. Do not repeat as fact.
+
+## Ruled out
+
+- Not the same as Northgate Mining Ltd (number 07654321). Different officers,
+ different address, name similarity only. Checked 2026-10-04.
+
+## Open
+
+- Beneficial owner behind the BVI parent. UK PSC filing names the parent only.
+```
+
+The `Ruled out` section is the one people skip and the one that saves the investigation. An
+explicitly rejected hypothesis stays rejected; an implicitly rejected one resurfaces in a month as
+a half-remembered lead and sometimes as a published error.
+
+Three cautions. Do not use wiki-style double-bracket links if the notes might ever be published or
+converted — they do not resolve outside the application, and a dead link in a published finding
+reads as sloppiness. Sync services put your entire research graph on third-party infrastructure, so
+if the vault syncs, know where to. And a vault is not an archive: the notes reference the artefacts,
+and the artefacts need their own hashes and timestamps.
+
+### The research account
+
+A separate identity that touches targets, which you expect to lose. Platforms ban research accounts
+eventually — for scraping, for viewing too many profiles, for geographic inconsistency, for nothing
+at all — so build nothing on it you cannot walk away from, and never let it share anything with an
+account you care about.
+
+There is no tool for this. There is a checklist, and the order matters:
+
+```text
+BEFORE the account exists
+ [ ] Decide the persona's purpose. A plausible-but-empty account is more
+ suspicious than an obviously new one with a stated interest.
+ [ ] Email first, from a provider that does not require a phone number, and
+ never from the provider your real accounts use.
+ [ ] Phone number, if the platform demands one. A number you control and can
+ lose. Never your own; a reused number links the research account to you
+ permanently and silently, because platforms match on it across services.
+ [ ] Password manager entry with the case ID in the title, so the account is
+ findable and disposable as a unit.
+
+WHEN the account exists
+ [ ] Separate browser profile at minimum, separate VM if the target is
+ competent. Never the same profile as anything personal — shared cookies,
+ shared localStorage and shared autofill all link them.
+ [ ] Consistent timezone, language and locale between the account's stated
+ location and the browser's. A profile claiming Lisbon from an en-GB
+ browser on UTC+0 is a detectable mismatch.
+ [ ] No contact import. Ever. One accidental contact sync hands the platform
+ your real address book and the platform hands the target "people you
+ may know".
+ [ ] Age the account before using it. A day-old account viewing 200 profiles
+ is a rate-limit and a ban; a two-month-old one is a user.
+
+ONGOING
+ [ ] One account per investigation where the targets could plausibly compare
+ notes. Cross-contamination between cases is how one burn becomes three.
+ [ ] Log every target the account touched, so you can assess the damage when
+ it is eventually identified.
+ [ ] Expect it to be lost. Export anything you need from it as you go.
+```
+
+The legal and ethical line is not a technical question and the checklist does not answer it.
+Creating an account usually breaches a platform's terms of service; in some jurisdictions and some
+employment contexts it does more than that, and a persona that actively deceives a person — rather
+than merely observing public content — is a different act from a passive research account. Know
+which one you are doing, get it authorised in writing if you are doing it for anyone but yourself,
+and note that nothing here makes impersonating a real person or organisation acceptable.
+
+### Isolation: a disposable VM
+
+A compromise or a deanonymisation should cost you a throwaway machine rather than your real one and
+the identity attached to it. Containers and virtual machines are not interchangeable here:
+a container shares the host kernel and, with default settings, the host network identity, which
+makes it fine for running CLI tooling and wrong for browsing a hostile target. Browse from a VM.
+
+```bash
+# full VM, Ubuntu, disposable: multipass on macOS and Windows
+multipass find # 26.04 is the current LTS alias
+multipass launch lts --name research-014 --cpus 2 --memory 4G --disk 20G
+multipass shell research-014
+multipass stop research-014 && multipass delete research-014 --purge # gone
+
+# or Lima, the same idea with a declarative YAML template.
+# Note the locator form: `template:` — the older `template://` is deprecated as of Lima 2.0
+limactl create --name=research-014 template:ubuntu-lts
+limactl start research-014
+limactl shell research-014
+limactl delete --force research-014
+
+# snapshot before you touch the target, so you can roll back to clean.
+# multipass only snapshots a STOPPED instance, so stop it first
+multipass stop research-014
+multipass snapshot research-014 --name pre-target
+multipass start research-014
+# ...after the session, roll back
+multipass stop research-014
+multipass restore research-014.pre-target --destructive
+```
+
+```bash
+# containers: correct for CLI tooling, with the network identity made explicit
+docker run --rm -it \
+ --dns 1.1.1.1 \
+ --cap-drop ALL --security-opt no-new-privileges \
+ -v "$PWD/out:/out" \
+ python:3.13-slim bash
+
+# route a container's traffic through a proxy you control, and nothing else
+docker run --rm -it \
+ -e ALL_PROXY=socks5h://host.docker.internal:1080 \
+ -e HTTPS_PROXY=socks5h://host.docker.internal:1080 \
+ -v "$PWD/out:/out" python:3.13-slim bash
+
+# a container that cannot reach the network at all, for handling a hostile file
+docker run --rm -it --network none -v "$PWD/sample:/sample:ro" python:3.13-slim bash
+
+# check the VM's egress before you use it, not after
+multipass exec research-014 -- curl -s https://am.i.mullvad.net/json
+```
+
+For network-level rather than machine-level isolation, Whonix routes an entire VM's traffic through
+Tor at the gateway, so a misconfigured application inside it cannot leak around the proxy, and
+Tails gives you an amnesic live system that forgets everything on shutdown. Both are the right
+answer when the consequence of being identified is serious; both are heavy enough that people skip
+them for routine work, which is a defensible trade as long as it is a decision rather than a
+default.
+
+What isolation does not buy you: a clean VM behind a VPN is still identified by the account you log
+into, the browser fingerprint you present and the timing of your activity. Isolation limits the
+blast radius of a mistake. It does not make you anonymous.
+
+### Touching a target without leaking
+
+When you need the bytes from a target's own server rather than from an archive, go via the command
+line rather than a browser, because a browser sends dozens of headers you did not choose and runs
+code the target wrote. Here the honest version matters: **neither `curl` nor `wget` has a
+`--no-referer` flag, because neither sends a `Referer` header unless you ask it to.** Verified, a
+default `curl` request arrives at the server carrying only `Host`, `User-Agent` and `Accept`.
+
+```bash
+# exactly what a default curl sends — check this yourself rather than trusting it
+curl -s https://postman-echo.com/headers | jq .
+# {"headers":{"host":"postman-echo.com","user-agent":"curl/8.7.1","accept":"*/*", ...}}
+```
+
+```bash
+# headers only, no body: cheapest possible look, and it reveals the stack
+curl -sI 'https://target.example/page'
+
+# response headers AND body, with the body discarded — some servers lie on HEAD
+curl -s -o /dev/null -D - 'https://target.example/page'
+
+# present a plausible browser UA. A default "curl/8.7.1" in a small site's log
+# is a flag that says "someone is scripting against us"
+curl -s -A 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0 Safari/537.36' \
+ 'https://target.example/page' -o page.html
+
+# send no User-Agent at all, which is a different and sometimes better choice
+curl -s -H 'User-Agent:' 'https://target.example/page' -o page.html
+
+# a Referer only when you deliberately want the log to show a plausible path
+curl -s -A 'Mozilla/5.0 ...' -e 'https://www.google.com/' 'https://target.example/page'
+
+# show the redirect chain without following it into something you did not expect
+curl -sIL -w '%{http_code} %{url_effective}\n' -o /dev/null 'https://target.example/short'
+
+# through a SOCKS proxy, with DNS resolved AT the proxy — socks5h, not socks5
+curl -s --proxy socks5h://127.0.0.1:9050 'https://target.example/page' -o page.html
+
+# wget equivalents, for a recursive pull you want to keep polite
+wget -U 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0 Safari/537.36' \
+ --wait=2 --random-wait --limit-rate=200k \
+ --page-requisites --adjust-extension 'https://target.example/page'
+```
+
+`socks5h` versus `socks5` is the one that bites people: with `socks5`, `curl` resolves the hostname
+locally and only the TCP connection goes through the proxy, so your resolver — and therefore your
+ISP and often your employer — sees exactly which host you looked up. With `socks5h` the proxy does
+the resolution. The same distinction applies to `ALL_PROXY` in the container examples above.
+
+Two things no flag fixes. The TLS handshake carries the hostname in SNI unless the server supports
+Encrypted Client Hello, so a passive observer on your network learns which site you contacted
+regardless of these options. And a target that fronts its site with a CDN sees your request through
+that CDN's logging, which is a second party you did not choose.
+
+### Leak checks: IP, DNS and WebRTC
+
+Run these before you touch a target, after every network change, and after every VPN reconnect.
+The failure you are looking for is not "am I behind a VPN" — it is "does something on this machine
+resolve or connect outside the tunnel", which is invisible until you measure it.
+
+```bash
+# the address a web server sees
+curl -s https://am.i.mullvad.net/json | jq '{ip, country, city, mullvad_exit_ip, organization}'
+curl -s https://ifconfig.co/json | jq '{ip, country, asn_org}'
+
+# the address your DNS RESOLVER egresses from — this is the leak that matters.
+# If this is your ISP while the line above is a VPN exit, DNS is outside the tunnel.
+dig +short TXT o-o.myaddr.l.google.com @ns1.google.com
+
+# resolver IP, your apparent client IP, and the EDNS Client Subnet your resolver
+# is handing to authoritative servers. An "ecs" line means your /24 is being
+# disclosed to every nameserver you query.
+dig +short TXT whoami.ds.akahelp.net
+
+# which resolvers this machine is actually using, whatever the VPN client claims
+scutil --dns | grep nameserver # macOS
+resolvectl status | grep -A2 'DNS Serv' # systemd-resolved
+
+# does a hostname resolve the same inside and outside the isolated VM
+dig +short target.example
+multipass exec research-014 -- dig +short target.example
+
+# mullvad's own CLI, if that is your provider: state, and whether DNS is leaking
+mullvad status
+mullvad dns get
+```
+
+WebRTC and browser fingerprinting have no CLI equivalent, because the leak is the browser's and
+only a browser reproduces it:
+
+```text
+https://browserleaks.com/webrtc does the browser disclose your real local
+ and public IP via ICE candidates, around
+ the proxy. This is the classic VPN leak and
+ it is a browser setting, not a network one.
+https://browserleaks.com/dns which resolvers the browser actually used
+https://www.dnsleaktest.com/ extended test; run it, not the standard one
+https://coveryourtracks.eff.org/ how distinctive your fingerprint is. Read
+ the "one in N browsers" number, not the
+ pass/fail badge.
+https://browserleaks.com/geo whether the page can get a precise location
+ from the OS rather than from the IP
+```
+
+Read the fingerprint result carefully, because the intuition is backwards: hardening a browser with
+unusual settings and a long extension list makes it *more* identifiable, not less. A stock browser
+in a stock VM is often the better disguise than a heavily customised one, and the only reliable
+counter to fingerprinting is to look like a large crowd.
+
+What these checks cannot tell you: whether the target correlated your visit with something else.
+Timing, a reused screen resolution, the same unusual font set, a session that starts every weekday
+at 09:10 UTC — none of that shows up in a leak test, and all of it is linkable across identities.
+
+### Chain of custody
+
+The record that lets you say, later, that the file you are producing is the file you received, and
+that nothing happened to it in between that you have not written down. It costs a minute per
+artefact and it is the first thing attacked when a finding matters.
+
+```bash
+# the moment an artefact arrives, before you open it
+IN=~/cases/2026-014/20-raw
+mkdir -p "$IN"
+cp /Volumes/USB/clip.mp4 "$IN/" # copy, never move
+shasum -a 256 "$IN/clip.mp4" | tee -a "$IN/MANIFEST.sha256"
+chmod 444 "$IN/clip.mp4" # read-only: work on copies
+
+# the custody note, as a sibling file, written now and never edited
+cat > "$IN/clip.mp4.custody" <<'EOF'
+artefact: clip.mp4
+sha256: 4b1c8e... # from MANIFEST.sha256
+received_at: 2026-10-04T04:07:56Z # date -u +%FT%TZ
+received_from: source S-3 (see 10-entities/source-s3.md), in person, USB
+provided_as: claimed original off a phone; no chain before this point
+handled_by: J. Investigator
+first_action: hashed, set read-only, copied to 50-work/ for analysis
+onward: none
+EOF
+
+# every derived file records what it came from, so no copy is ever orphaned
+shasum -a 256 50-work/clip-frame-0137.png >> 50-work/DERIVED.sha256
+printf '%s\tderived from\t%s\n' 'clip-frame-0137.png' 'clip.mp4 (4b1c8e...)' \
+ >> 50-work/DERIVED.index
+
+# append-only activity log, one line per action, in UTC
+printf '%s\t%s\n' "$(date -u +%FT%TZ)" 'extracted frame 00:01:37.5 with ffmpeg 8.0' \
+ >> ~/cases/2026-014/90-log.md
+
+# verify the whole raw tree before you hand anything over
+shasum -a 256 -c "$IN/MANIFEST.sha256" | grep -v ': OK$'
+```
+
+Then timestamp the manifest, which is what converts your own record into a third party's assertion
+about time — `ots stamp` and the RFC 3161 route are in
+[Archiving & Evidence Preservation](/sheets/osint/archiving-and-evidence).
+
+Four rules that are not negotiable. UTC everywhere, because a local timestamp in a report with
+international sources is ambiguous and ambiguity is attackable. Copy rather than move, so the
+original stays where it was. Write the note at the time, because a custody note reconstructed from
+memory is exactly as reliable as it sounds. And record the gap honestly: if you do not know where
+the file was before your source handed it to you, the note says "no chain before this point" rather
+than nothing, because an unstated gap reads as a concealed one.
+
+### The negative-results log
+
+The cheapest high-value habit on this page, and the one almost nobody keeps. A searchable record of
+every query that returned nothing stops you repeating dead ends, tells you where your coverage
+actually ends, and is the only honest basis for a sentence like "no public record of X exists".
+
+```text
+# 40-negative/not-found.md — append-only, one block per attempt
+
+## 2026-10-04T04:12Z — Companies House officer search
+query: "Khan" + date of birth 1979-03
+scope: all UK registered companies, active and dissolved
+result: 0 matches with that DOB month
+means: not registered as a UK officer under that spelling. Does NOT mean
+ not an officer: Companies House shows DOB month and year only, and
+ transliteration variants were not tested.
+next: re-run via Bellingcat Name Variant Search for Cyrillic variants
+
+## 2026-10-04T04:31Z — Sherlock, handle "acct_f"
+query: sherlock acct_f --print-found
+result: 3 hits, all confirmed unrelated third parties
+means: handle is reused; it is not a usable pivot for this person
+next: pivot on the profile photo instead, not the handle
+
+## 2026-10-04T05:02Z — Wayback, target.example/about
+query: CDX, matchType=prefix, from=20240101
+result: no captures
+means: nothing about absence of the page. The site serves a robots
+ directive that Wayback honoured at crawl time.
+next: check archive.today and Ghostarchive before concluding anything
+```
+
+The `means:` line is the whole point and it is the line that gets dropped. "Zero results" is a fact
+about a tool's coverage, and turning it into a fact about the world is the single most common
+overreach in open-source work. A people-search service returning nothing means that service has
+nothing. A registry returning nothing means that registry, under that spelling, on that date.
+
+Write the `next:` line too. A negative result with a stated next step is a lead; a negative result
+without one is just a gap you will rediscover.
## Verification
@@ -94,6 +535,125 @@ Two sources that both trace back to the same original post are one source.
use it.
- **Forgetting the human cost.** Publishing that someone can be located has consequences for them.
Minimise what you expose beyond what the finding requires.
+- **A VPN indicator is not a leak test.** The client says "connected" while the system resolver
+ still egresses from your ISP, and a browser can hand out your real address over WebRTC around any
+ tunnel. Measure the resolver and the browser separately, after every reconnect.
+- **`socks5` where you meant `socks5h`.** The proxy carries the connection but your own resolver
+ does the lookup, so the hostname is logged locally even though the traffic was not.
+- **Hardening a browser makes it more identifiable.** An unusual configuration and a long extension
+ list are a fingerprint. A stock browser in a disposable VM hides in a larger crowd.
+- **Zero results is a fact about a tool, not about the world.** Record what the gap means and what
+ it does not, in the same note, at the time.
+
+## Worked example
+
+One datum: a tip-off email naming a company, `Northgate Minerals Trading`, and asserting it is a
+front. Nothing else. The question is not "is it a front" — that is a conclusion looking for support.
+This is the first hour, before any collection, and what it buys you.
+
+```text
+# 1. write the question down. 00-question.md
+question: Does any public record link Northgate Minerals Trading to
+ [named entity] through a shared officer, address or owner?
+standard: two independent primary sources per link, or it is "reported, unverified"
+out of scope: the tipster's motive; family members of any officer
+stop when: the question is answered either way, or three named registries
+ have been checked and returned nothing
+exposure: registry APIs and archives only. NO requests to any Northgate-
+ controlled domain from an attributable address until step 6.
+```
+
+The exposure line is the decision that constrains everything after it. Having written it down, the
+first four steps are forced: only sources that do not touch the target.
+
+```bash
+# 2. establish the egress before anything leaves the machine
+curl -s https://am.i.mullvad.net/json | jq '{ip, country, organization, mullvad_exit_ip}'
+dig +short TXT o-o.myaddr.l.google.com @ns1.google.com
+# web-visible IP: a VPN exit, country NL
+# resolver egress: 203.0.113.x — the SAME ISP range as the host, not the VPN
+```
+
+DNS is outside the tunnel. Every hostname you look up is visible to the ISP, and will be whether or
+not the HTTP request goes through the VPN. Fix that before step 3, not after: the resolver leak is
+the one that persists, because it is a system setting and the VPN client reported "connected".
+
+```bash
+# 3. isolation, and verify it from inside rather than trusting the launch
+multipass launch lts --name research-014 --cpus 2 --memory 4G --disk 20G
+multipass exec research-014 -- curl -s https://am.i.mullvad.net/json | jq -r .ip
+multipass exec research-014 -- dig +short TXT o-o.myaddr.l.google.com @ns1.google.com
+# both now report the VPN exit. Snapshot clean before use:
+multipass stop research-014 && multipass snapshot research-014 --name pre-target
+multipass start research-014
+```
+
+```text
+# 4. open the Hunchly case BEFORE the first search, not after
+case name: 2026-014 (the case ID, never "Northgate")
+selectors: Northgate Minerals Trading
+ Northgate Ltd
+ 09876543
+ [officer surname, once known]
+capture: ON — confirmed by the extension icon after the browser restart
+```
+
+Every page from here on is captured, hashed and timestamped whether or not you thought it mattered
+at the time. That is the whole reason the case opens before the searching: the page you will need is
+the one you skimmed on the way to something else.
+
+```bash
+# 5. collect from sources that do not touch the target, and log the negatives
+# (the registry and sanctions commands themselves live on the companies sheet)
+printf '## %s — Companies House name search\nquery: "Northgate Minerals Trading"\nresult: 2 hits, gb/09876543 active + vg/1654321 (source dated 2019)\nmeans: a UK and a BVI company share the name. Not yet evidence they are related.\nnext: PSC filing on 09876543\n\n' \
+ "$(date -u +%FT%TZ)" >> ~/cases/2026-014/40-negative/not-found.md
+```
+
+```bash
+# 6. the only step that touches the target, and only after deciding to
+# headers first: it answers the hosting question without fetching the page
+multipass exec research-014 -- \
+ curl -sI -A 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0 Safari/537.36' \
+ 'https://northgate-minerals.example/' | head -12
+# Server: cloudflare → the origin IP is not in this response, and a CDN
+# operator now has a log line for this request. Noted in 90-log.md.
+```
+
+A `HEAD` from a VM behind a VPN, with a browser UA, after an explicit decision — rather than a
+browser tab opened reflexively in hour one from your own address. The difference costs ninety
+seconds and is the difference between the target knowing and not knowing.
+
+```bash
+# 7. the artefacts, hashed and custody-noted at the moment they arrive
+IN=~/cases/2026-014/20-raw; mkdir -p "$IN"
+multipass transfer research-014:/home/ubuntu/ch-09876543.json "$IN/"
+shasum -a 256 "$IN/ch-09876543.json" | tee -a "$IN/MANIFEST.sha256"
+chmod 444 "$IN/ch-09876543.json"
+printf '%s\t%s\n' "$(date -u +%FT%TZ)" 'retrieved CH filing 09876543 via API from research-014' \
+ >> ~/cases/2026-014/90-log.md
+```
+
+```text
+# 8. the entity note, written as three separate claims. 10-entities/company-northgate.md
+Established: number 09876543, UK, active. Source: CH API, 2026-10-04T05:12Z,
+ raw ch-09876543.json (sha256 4b1c8e...).
+Reported/unverified: "a front". Single source, the tipster, who has an interest.
+ Not repeated as fact anywhere else in the case.
+Ruled out: not Northgate Mining Ltd (07654321) — different officers,
+ different address, name similarity only. Checked 2026-10-04.
+Open: beneficial owner behind the BVI parent.
+```
+
+After an hour the case holds: a written question with a stop condition, a fixed and verified
+exposure posture, a continuous capture log, two registry records with hashes and custody notes, one
+explicitly rejected lookalike company, and a negative-results entry that says what "two hits" does
+and does not mean.
+
+What none of that establishes: whether Northgate is a front. The tipster's claim is still exactly
+one uncorroborated assertion, recorded as such, in a section of the note that cannot be mistaken
+for a finding. The substantive work now moves to
+[Company & Financial Records](/sheets/osint/companies-and-finance) — but it moves there on top of a
+record that will survive someone attacking it, which is the only thing this sheet is for.
## Broader catalogues
diff --git a/src/content/sheets/osint/reverse-image-search.md b/src/content/sheets/osint/reverse-image-search.md
@@ -4,9 +4,9 @@ description: "Find where an image came from and whether it predates the event it
category: osint
subcategory: "Images & Video"
tags: [osint, images, verification, reverse-search]
-tools: [google-lens, tineye, yandex, invid]
+tools: [google-lens, tineye, yandex, bing, invid, ffmpeg, imagemagick]
difficulty: beginner
-updated: 2026-09-28
+updated: 2026-10-04
references:
- name: "Bellingcat's Online Investigation Toolkit"
url: "https://bellingcat.gitbook.io/toolkit"
@@ -24,72 +24,361 @@ references:
## What this covers
-Establishing whether an image is what it claims to be, by finding earlier copies. This is the
+Establishing whether an image is what it claims to be, by finding earlier copies of it. This is the
single highest-value check in visual verification and it takes under a minute, so it goes first,
-always. Most viral misinformation is old footage relabelled.
+always — most viral misinformation is old footage relabelled, and the relabelling is what a reverse
+search catches. The engines disagree constantly, which is not a flaw to work around but the reason
+you run more than one.
## Method
-1. **Run several engines.** They index different corpora and disagree constantly. One engine
- returning nothing means nothing.
-2. **Sort by oldest**, not by relevance. You want the earliest appearance, which is the likely
- original.
-3. **Crop and re-run.** Engines match on the whole frame. Cropping to a distinctive element — a
- sign, a building, a vehicle marking — often finds matches the full frame misses.
-4. **For video, search keyframes.** Extract frames and search each; a video is only as findable as
- its most distinctive still.
-5. **Check the earliest hit's own context.** The oldest copy you can find is not necessarily the
- original — read its caption and follow its sourcing.
+1. **Run several engines on the unmodified file.** They index different corpora. One engine
+ returning nothing means that engine has nothing, and nothing more than that.
+2. **Sort by oldest where the engine lets you.** Relevance ranking is actively unhelpful here —
+ you want the earliest appearance, not the most popular one.
+3. **Preprocess and re-run.** Crop to the distinctive object, remove the overlay, flip it
+ horizontally, upscale it. Each of these is a different query against the same corpus and each
+ finds things the others miss.
+4. **For video, extract keyframes and search each.** A video is only as findable as its most
+ distinctive still.
+5. **Read the earliest hit's own context.** The oldest copy you can find is not necessarily the
+ original. Follow its caption and its sourcing before you call it the source.
+6. **Archive what you found, with its date, before you cite it.** Search results are not stable —
+ see [Archiving & Evidence](/sheets/osint/archiving-and-evidence).
-## Engines, and what each is good for
+The judgement call that matters: decide what you are actually asking. "Where did this image come
+from" and "does this image predate the event it is captioned with" are different questions, and
+only the second one is usually answerable. The second is also the one that settles arguments, so
+aim at it.
-| Engine | Strength |
-| --- | --- |
-| [Google Lens](https://lens.google.com/) | Best at objects, text in images, landmarks and products. Strong OCR. |
-| Yandex Images | Consistently the best at faces and at Eastern European and Central Asian content. Frequently finds what Google misses. |
-| [TinEye](https://tineye.com/) | Oldest-first sorting and exact-match focus. The best tool for "when did this first appear". |
-| Bing Visual Search | Good regional coverage; worth running as a third opinion. |
-| [Search by Image](https://addons.mozilla.org/en-US/firefox/addon/search_by_image/) | Browser extension that fires one image at many engines at once. |
-| [RootAbout](https://rootabout.com/) | Reverse image search across Internet Archive holdings. |
+## Key tools
+
+### Google Lens
+
+Web only, free, no account needed for a one-off. The strongest engine for *objects* — products,
+landmarks, vehicles, plants, architecture — and it has the best OCR of any of them, which means it
+will often read a sign in your image and search the text without being asked. That is exactly what
+you want for a photo with a shopfront in it, and exactly what you do not want for a photo whose
+text is incidental.
+
+```text
+1. lens.google.com, or the camera icon in Google Images. Upload the file; do not
+ paste a social-media URL, which searches the page rather than the image.
+2. Drag the crop handles immediately. Lens defaults to the whole frame and the whole
+ frame is almost always the wrong query. Crop to one object and re-run.
+3. Switch between the result modes. "Exact matches" is the one that answers the
+ verification question; the visual-similarity feed is for identifying the object.
+4. "Find image source" (where offered) gives a dated list of pages carrying the image.
+ This is the closest Lens gets to TinEye's oldest-first view, and it is worth
+ scanning to the bottom rather than the top.
+5. Capture: the result page as a screenshot plus the URLs of the earliest pages, and
+ the crop you actually searched -- a result is not reproducible without the crop.
+```
+
+You can also hand it a URL directly, which is useful for scripting a triage pass over a set of
+images already hosted somewhere:
+
+```bash
+# searches the image at that URL; the image must be publicly reachable
+open "https://lens.google.com/uploadbyurl?url=https://example.org/photo.jpg"
+
+# url-encode anything with query parameters or the search will be of the wrong file
+python3 -c 'import sys,urllib.parse; print("https://lens.google.com/uploadbyurl?url=" + urllib.parse.quote(sys.argv[1], safe=""))' \
+ "https://example.org/img?id=123&size=full"
+```
+
+Lens personalises, so two people searching the same image get different result sets and neither is
+reproducible from the other's screenshot. It is also weak on faces by design — it will refuse or
+deflect — and weak on anything that is mostly sky, water or crowd. Dates shown next to results are
+the page's claimed date, not the image's, and content farms backdate.
+
+### Yandex Images
+
+Web only, free. The one that finds what Google misses, consistently enough that skipping it is a
+real gap in a verification. It is markedly better at faces, at crops of faces, and at Eastern
+European, Russian and Central Asian content of every kind — if the image has Cyrillic in it or is
+plausibly from that part of the world, Yandex is your first engine rather than your third.
+
+```text
+1. yandex.com/images -> the camera icon -> upload, or paste an image URL.
+2. The results page splits into "Similar images" and "Sites containing this image".
+ The second list is the evidential one; the first is a lookalike feed and will
+ cheerfully show you a different person with the same haircut.
+3. Use the size filter in the left rail. Filtering to larger-than-your-copy is a fast
+ way to find an earlier, less-compressed version, which is usually closer to source.
+4. Re-run on a tight crop of a face or a patch, separately. Yandex handles partial
+ matches far better than the others and a crop frequently returns hits the full
+ frame does not.
+5. Read the result page titles even when the thumbnails look wrong -- its indexing of
+ Russian-language forums and VK is better than anything else available.
+```
+
+The same URL pattern works if you want to fire it from a script:
+
+```bash
+open "https://yandex.com/images/search?rpt=imageview&url=https://example.org/photo.jpg"
+```
+
+The face matching is good enough to be an ethical question rather than a technical one: it will
+identify private individuals from a crowd shot. Decide before you run it what you will do with a
+hit, and whether the person in the frame is a subject of your investigation or a bystander. Beyond
+that, Yandex's similarity feed is the most confident and the most wrong of any engine — it will
+return a visually similar image with total assurance, so never treat a "similar" result as a match
+without comparing fixed detail yourself.
+
+### TinEye
+
+Web only. Free for interactive use with no account; the published free API tier is **100 searches
+per day and 300 per week** for non-commercial use, with paid tiers above that. The index is smaller
+than Google's and it does not do visual similarity at all — it does exact and near-exact matching
+of the same image, including crops, resizes and recolours.
+
+That narrowness is the whole point: **TinEye's "Oldest" sort is the single strongest piece of
+evidence available for "this image predates the event it claims to show."** If TinEye returns a
+copy indexed in 2019 and the caption says the photo is from last week's protest, the caption is
+false, and no amount of argument about context changes that.
+
+```text
+1. tineye.com -> upload the file, or paste an image URL.
+2. Change the sort from "Best Match" to "Oldest". This is the only control on the page
+ that matters for verification. The options are Best Match, Most Changed, Biggest
+ Image, Newest and Oldest.
+3. Read the first result's crawl date and open the page it sits on. TinEye's date is
+ when it first saw the image at that URL, not when the image was made -- it is a
+ latest-possible-creation date, which is the useful direction.
+4. "Most Changed" is the second-most-useful sort: it surfaces the versions that have
+ been cropped, overlaid or edited, which shows you how the image has been used.
+5. "Biggest Image" finds the least-compressed copy, which is what you want to hand to
+ the other engines and to any forensic step.
+6. Capture: the oldest result's URL, TinEye's stated date for it, the total match
+ count, and an archive snapshot of that oldest page before it moves.
+```
+
+Absence from TinEye is close to meaningless — its crawl is much narrower than Google's and it does
+not index most social platforms, so a genuinely viral image can show zero results. The date is a
+crawl date and a page can have been republished at a new URL, so an "oldest" of last month does not
+mean the image is from last month. And it is defeated by heavy re-editing in a way the
+similarity-based engines are not: a mirrored, re-captioned, re-encoded repost may not match at all,
+which is why the preprocessing step below exists.
+
+### Bing Visual Search
+
+Web only now. The Bing Search APIs — including Visual Search — were **retired on 11 August 2025**
+and pre-retirement endpoints return HTTP 410, so any tutorial or script that calls a Bing visual
+search API is dead; Microsoft's replacement is a grounding feature inside Azure AI Agents rather
+than an image-match endpoint. The web interface still works and is still worth a minute, because
+its regional coverage differs from both Google's and Yandex's.
+
+```text
+1. bing.com/images -> the camera icon in the search box. (bing.com/visualsearch now
+ redirects to a Microsoft marketing page, not the tool.)
+2. Upload, paste a URL, or drag a file in.
+3. Use the on-image crop box, which Bing exposes more prominently than Google does --
+ it is the best of the three for "search just this one object in the picture".
+4. "Pages with this image" is the evidential list; "Related content" is not.
+5. Capture the same way as the others: screenshot, URLs, and the crop you used.
+```
+
+Scriptable as a URL, with the usual caveat that it is an undocumented web parameter rather than a
+supported interface and may change without notice:
+
+```bash
+open "https://www.bing.com/images/search?view=detailv2&iss=sbi&q=imgurl:https://example.org/photo.jpg"
+```
-Yandex's face matching is good enough to be an ethical question, not just a technical one. Think
-about what happens to the person in the photo if you identify them.
+No oldest-first sort, no date on most results, and a strong pull toward commercial and stock
+imagery. Treat it as a third opinion that occasionally produces the one hit nobody else had,
+rather than as a primary tool.
-## Video keyframes
+### Preprocessing with ImageMagick
-[InVID / WeVerify](https://www.invid-project.eu/tools-and-services/invid-verification-plugin/) is
-the standard browser plugin: it extracts keyframes from a video, runs them through several reverse
-image engines, and also surfaces upload metadata and a magnifier for detail work.
+The step that separates a reverse search that works from one that returns nothing. Engines match on
+what is visually dominant, so an overlaid caption, a platform watermark or a wide establishing shot
+all push the match toward the wrong thing. Each transformation below is a *separate query* — run
+them all, do not pick one.
-Manually, with ffmpeg:
+```bash
+brew install imagemagick # or: apt install imagemagick
+
+# what you are actually working with: dimensions tell you how much has been lost already
+magick identify -verbose photo.jpg | head -40
+
+# crop to the distinctive object. percentages, or pixels as WxH+X+Y
+magick photo.jpg -crop 40%x40%+30%+20% +repage crop-sign.jpg
+magick photo.jpg -crop 640x480+120+80 +repage crop-vehicle.jpg
+
+# mirror it. reposts are flipped constantly to defeat exact matching, and the engines
+# do not try both orientations for you
+magick photo.jpg -flop flipped.jpg
+
+# cut off a burned-in caption bar or platform watermark along the bottom
+magick photo.jpg -gravity south -chop 0x90 nobar.jpg
+
+# upscale a small crop so an engine has pixels to work with. lanczos then a light
+# unsharp is the combination that does not invent edges
+magick crop-sign.jpg -filter Lanczos -resize 300% -unsharp 0x1 crop-up.jpg
+
+# pull detail out of a dark or flat region before cropping to it
+magick photo.jpg -auto-level -sigmoidal-contrast 3,50% enhanced.jpg
+
+# strip metadata from anything you are about to upload to a third-party engine --
+# you are handing them the file, and the file may carry a source's GPS
+magick photo.jpg -strip clean.jpg
+```
+
+`+repage` after a crop is not optional: without it the file keeps the original canvas geometry and
+some tools will re-expand it. Upscaling adds no information — it makes a small crop palatable to an
+engine's minimum-size requirements, and anything it appears to reveal is interpolation, so never
+read a plate number off an upscale. And every enhancement you apply is a change to evidence: keep
+the original untouched, work on copies, and record the exact command alongside the result, the same
+way [Image & Video Forensics](/sheets/osint/image-video-forensics) handles hashing and the chain
+from received file to analysed file.
+
+### ffmpeg
+
+Video does not reverse-search. Frames do. Pull the frames that are worth searching and you have
+turned an unsearchable clip into eight or ten image queries.
```bash
-# scene-change keyframes — the frames worth searching
-ffmpeg -i video.mp4 -vf "select='gt(scene,0.3)'" -vsync vfr keyframe-%03d.jpg
+brew install ffmpeg # or: apt install ffmpeg
+
+# every encoded keyframe, decoded fast because P- and B-frames are skipped outright.
+# this is the cheapest first pass and usually enough
+ffmpeg -skip_frame nokey -i video.mp4 -fps_mode vfr -q:v 2 key-%03d.jpg
+
+# frames at scene changes: 0.3 is a sane threshold, lower it to 0.1 for static footage
+ffmpeg -i video.mp4 -vf "select='gt(scene,0.3)'" -fps_mode vfr scene-%03d.jpg
+
+# one frame every five seconds, as a fallback for a single unbroken shot that has no
+# scene changes to detect
+ffmpeg -i video.mp4 -vf fps=1/5 every5-%03d.jpg
+
+# a contact sheet, to pick the searchable frame by eye instead of opening fifty files
+ffmpeg -i video.mp4 -vf "select='gt(scene,0.1)',scale=320:-1,tile=4x4" -fps_mode vfr sheet.png
+
+# the exact frame at a timestamp, once you know which second carries the readable sign
+ffmpeg -ss 00:01:12.500 -i video.mp4 -frames:v 1 frame.png
-# one frame every five seconds, as a fallback
-ffmpeg -i video.mp4 -vf fps=1/5 frame-%03d.jpg
+# crop, upscale and sharpen that frame in one pass, ready to hand to an engine
+ffmpeg -i frame.png -vf "crop=400:300:820:460,scale=iw*4:ih*4:flags=lanczos,unsharp=5:5:1.0" search-me.png
-# a contact sheet for eyeballing the whole video at once
-ffmpeg -i video.mp4 -vf "select='gt(scene,0.2)',scale=320:-1,tile=4x4" -vsync vfr sheet.png
+# mirror a frame, for the flipped-repost case
+ffmpeg -i frame.png -vf hflip frame-flipped.png
+
+# blank out a burned-in platform logo rather than cropping the frame down
+ffmpeg -i frame.png -vf "delogo=x=20:y=20:w=160:h=60" nologo.png
```
+`-vsync vfr` appears in most tutorials for this and has been **removed** in ffmpeg 9 — it now fails
+with "Option not found". Use `-fps_mode vfr`, which replaced it in ffmpeg 5. Scene detection also
+fires on camera pans and on cuts to black as readily as on a genuine change of location, so expect
+a third of the frames to be useless; conversely a single-shot clip yields no scene frames at all,
+which is why the `fps=1/5` fallback is there. Downloading the video in the first place, and reading
+its container metadata, are on [Social Media Platforms](/sheets/osint/social-media-platforms) and
+[Image & Video Forensics](/sheets/osint/image-video-forensics) respectively.
+
+### InVID / WeVerify
+
+A browser extension, free, from an EU research project. It collapses the whole video path above
+into a few clicks and — the part that genuinely saves time — it extracts keyframes from a platform
+URL without you downloading the file at all, then fires each keyframe at several engines in one go.
+
+```text
+1. Install from weverify.eu/verification-plugin (Chrome and Firefox builds).
+2. "Keyframes": paste the video URL. It fetches, segments and returns a grid of
+ keyframes with reverse-search buttons under each one.
+3. Click through Google / Yandex / Bing / TinEye per keyframe. The buttons open the
+ engines in new tabs -- you still read the results yourself, it just saves the
+ upload step for each of the four.
+4. "Magnifier": loads a still into a zoom-and-enhance panel, for reading a sign
+ without leaving the browser.
+5. "Analysis": pulls the platform's own upload metadata for a YouTube/Facebook/X URL,
+ which gives you a latest-possible date for the upload -- not for the footage.
+6. "Image forensics" runs ELA and noise filters. Use it for triage only; the real
+ version of that work is on the forensics sheet.
+```
+
+Platform integrations break whenever a platform changes its markup, so expect at least one tab of
+the plugin to be dead at any given time; the keyframe extraction and the reverse-search buttons are
+the durable parts. It gives you no reproducible command and no file hash, so anything you intend to
+publish should be re-done with `ffmpeg` and recorded properly. And it uploads frames to third-party
+engines, which is the same exposure question as any web tool — for sensitive material, extract
+locally and decide deliberately what leaves your machine.
+
+### The smaller engines
+
+Worth a minute each when the big four come back empty, because they index corpora the others do not
+touch.
+
+| Tool | What it is good for |
+| --- | --- |
+| [RootAbout](https://rootabout.com/) | Reverse image search across Internet Archive holdings — old web, scanned books and ephemera that no live-web crawler has. |
+| [Search by Image](https://addons.mozilla.org/en-US/firefox/addon/search_by_image/) | Browser extension that fires one image at a configurable list of engines at once. The fastest way to run the full set without uploading four times. |
+| [Karma Decay](http://karmadecay.com/) | Reverse search restricted to Reddit, which is where a surprising share of recycled images surfaces first. |
+| [Pimeyes](https://pimeyes.com/) | Face search specifically, and paid beyond a teaser. Powerful, legally fraught in several jurisdictions, and a tool to think hard about before using on anyone who is not a subject. |
+
+None of these replaces the main four. Reaching for them is a sign that your query is wrong more
+often than it is a sign that the image is unindexed — go back and crop harder before you go
+further down this list.
+
## Tool reference
| Tool | What it does | Cost |
| --- | --- | --- |
| [InVID](https://weverify.eu/verification-plugin/) | A toolkit that supports the verification of videos and images. | free |
+| [Google Lens](https://lens.google.com/) | Object, landmark and text matching with strong OCR. | free |
+| [TinEye](https://tineye.com/) | Exact and near-exact matching with oldest-first sorting. | free / paid API |
+| [RootAbout](https://rootabout.com/) | Reverse image search across Internet Archive holdings. | free |
## Pitfalls
-- **Cropped, mirrored or filtered images defeat exact matching.** Flip the image horizontally and
- re-run; recompressed and mirrored reposts are extremely common.
-- **Absence of results is not originality.** It often just means the original is on a platform the
- engine cannot index.
-- **The oldest hit can still be a repost.** Read its context rather than treating the date as the
- answer.
-- **Screenshots of screenshots** lose the detail engines match on. Ask for the original file when
- you can.
+- **Cropped, mirrored or filtered images defeat exact matching.** Flip horizontally and re-run;
+ recompressed and mirrored reposts are extremely common and TinEye in particular will miss them.
+- **Absence of results is not originality.** It usually means the original lives on a platform the
+ engine cannot index. Say "not found by X, Y and Z", never "original".
+- **The oldest hit can still be a repost.** Read its context rather than treating the crawl date as
+ the answer.
+- **Dates on result pages are the page's claim.** Content farms backdate, CMSs rewrite timestamps
+ on edit, and an archive snapshot is the only date you can stand behind.
+- **Screenshots of screenshots** lose the detail engines match on. Ask for the original file.
+- **Results are personalised and are not reproducible.** Archive the result page, and always record
+ the crop you searched — without it, nobody can repeat your query.
+- **Uploading is disclosure.** Every web engine on this page keeps what you give it. For a file
+ from a source, strip it first and think about whether it should be uploaded at all.
+
+## Worked example
+
+One datum: a photograph circulating with the claim that it shows a named street during a protest
+three days ago.
+
+1. **Unmodified file, four engines.** Lens returns stock-photo lookalikes. Bing returns news
+ aggregators from this week. Yandex returns a Russian-language forum thread. TinEye, sorted
+ **Oldest**, returns a first crawl dated **four years ago** on a regional news site.
+2. **Open the oldest page.** It carries the same image, uncropped, with a caption naming a
+ different city and a different event. The claim is already broken at this point, and everything
+ after this is confirming rather than discovering.
+3. **Confirm it is the same image, not a lookalike.** Crop both copies to the same corner — a shop
+ awning with a readable name — and compare. Identical awning, identical crack in the paving,
+ identical parked van. Fixed detail agreeing is the test; general resemblance is not.
+4. **Explain why the other engines missed it.** The circulating version is mirrored and has a
+ caption bar burned along the bottom. `magick circulating.jpg -flop -gravity south -chop 0x90
+ fixed.jpg` undoes both, and a re-run gets Lens to the same original — which demonstrates the
+ edit and shows the preprocessing step earning its place.
+5. **Date the circulating version, not just the original.** TinEye's "Newest" sort and the
+ aggregator pages put the relabelled version's first appearance at four days ago, a day before
+ the protest it is captioned with — which is itself a finding.
+6. **Archive before citing.** The four-year-old page, the aggregator pages, and the TinEye result
+ page all go to the Wayback Machine, because the regional news site is exactly the kind of
+ source that reorganises its URLs. See [Archiving & Evidence](/sheets/osint/archiving-and-evidence).
+
+What you can assert: the image was indexed on a named site four years before the event it is
+captioned with, the circulating copy is a mirrored and cropped derivative of it, and specific fixed
+details match between the two.
+
+What would falsify it: the two images being different photographs of the same unchanged street —
+which is what the awning crack and the parked van rule out, and which is why you compare fixed
+detail rather than overall appearance. A crawl date that TinEye got wrong would also do it, so the
+archived copy of the four-year-old page, with its own publication date, is the thing worth having.
## Broader catalogues
diff --git a/src/content/sheets/osint/social-media-monitoring.md b/src/content/sheets/osint/social-media-monitoring.md
@@ -4,9 +4,9 @@ description: "Track accounts, hashtags and narratives across several platforms a
category: osint
subcategory: "Social Media"
tags: [osint, monitoring, collection, social-media]
-tools: [4cat, distill, snscrape]
+tools: [4cat, zeeschuimer, changedetection.io, distill, yt-dlp, rsshub, rss-bridge, arctic-shift]
difficulty: intermediate
-updated: 2026-09-28
+updated: 2026-10-04
references:
- name: "Bellingcat's Online Investigation Toolkit"
url: "https://bellingcat.gitbook.io/toolkit"
@@ -24,42 +24,444 @@ references:
## What this covers
-Working several platforms at once, and watching things over time rather than looking once. Two
-distinct jobs: **bulk collection** for later analysis, and **change detection** on pages you care
-about.
+Watching things over time instead of looking once, and holding what you collect so you can analyse
+it later without re-collecting. Two distinct jobs with different tooling: **bulk collection** of a
+corpus for analysis, and **change detection** on a small number of pages you care about. The third
+thing, which is not optional, is doing either without the pattern of your collection becoming
+visible to the people you are collecting on.
-## Bulk collection
+## Method
-[4CAT](https://github.com/digitalmethodsinitiative/4cat) is the serious option — a self-hosted
-capture-and-analysis platform that ingests from many platforms and ships analytical modules
-(co-word networks, time series, image clustering) on top of what it collects.
+1. **Decide what you are watching before you build anything.** A corpus for network analysis and an
+ alert on one bio edit need completely different infrastructure, and building the first when you
+ needed the second is the usual waste.
+2. **Prefer the platform's own feed.** A native RSS or JSON endpoint is stable, keyless, polite and
+ does not attribute activity to an account. Reach for a scraper only when no feed exists.
+3. **Store raw, analyse later.** Keep the untouched response alongside anything you derive from it.
+ You will want to re-run the analysis with a different question, and you will not want to
+ re-collect under a worse rate limit.
+4. **Record gaps as gaps.** A failed poll is a hole in your data, not a quiet day. Log the failures
+ next to the results or your time series will lie to you.
+5. **Set the cadence to the question.** Hourly is almost never justified. A daily poll that runs
+ for a year is more valuable, and far less conspicuous, than a five-minute poll that gets you
+ blocked in a week.
+6. **Baseline before you conclude.** Coordination, surges and silences only mean something against
+ what normal looks like for that topic.
+
+The judgement calls: whether you need the content or only the fact that it changed — the second is
+far cheaper and far quieter. Whether the collection needs an account at all, because the moment it
+does, the collection has an identity and a history. And whether you are allowed to keep what you
+are about to collect, which is a question to settle at the start rather than after you have a
+database of it.
+
+## Collection hygiene
+
+Monitoring is repeated, scheduled, patterned contact with a target's content. One look is
+invisible. The same request every fifteen minutes from one address for six months is a signature,
+and on some platforms it is a signature attached to a logged-in account.
+
+```text
+Account
+ - never your own, and never one that shares a recovery phone or email with your own
+ - a research account per investigation where the platform allows it; one burned
+ account should not cost you the others
+ - assume every logged-in query is retained and attributable, including search terms
+ - some platforms notify a user when a profile is viewed; know which before you look
+
+Rate
+ - daily is the default. Justify anything faster to yourself in writing
+ - randomise the interval. A poll at exactly :00 every hour is machine-obvious
+ - respect the stated limit, and treat a 429 as a signal to back off for hours,
+ not to retry in a loop
+ - set a real, honest User-Agent that identifies the project. It gets you unblocked
+ more often than a spoofed browser string does
+
+Network
+ - one address for all of it correlates every collection you run
+ - a residential or VPN exit is a trade: less rate limiting, more attribution risk
+ to whoever pays for it
+ - Tor is blocked by most of these platforms, so it is rarely the answer here
+
+Footprint
+ - log every request you make, with timestamp and response code. You need it to
+ distinguish "they stopped posting" from "we got blocked"
+ - decide retention up front. Bulk personal data attracts obligations even when
+ every item was public
+```
+
+The archiving and provenance side of this — hashes, snapshots, what makes a collected item citable
+later — is on [Archiving & Evidence](/sheets/osint/archiving-and-evidence). Per-platform surfaces,
+query syntax and the specific scrapers are on
+[Social Media Platforms](/sheets/osint/social-media-platforms); this sheet is about running those
+things repeatedly rather than once.
+
+## Key tools
+
+### 4CAT
+
+A self-hosted capture-and-analysis platform. It is the serious option because collection and
+analysis live in the same place: the dataset you capture stays on your machine, and you can re-run
+a different analysis over it months later without touching the platform again. That is the thing
+no hosted service gives you.
```bash
-git clone https://github.com/digitalmethodsinitiative/4cat
-cd 4cat && docker compose up -d
-# web UI is then served on port 80
+# the documented install is the compose file plus its .env, not a clone
+mkdir 4cat && cd 4cat
+curl -O https://raw.githubusercontent.com/digitalmethodsinitiative/4cat/master/docker-compose.yml
+curl -O https://raw.githubusercontent.com/digitalmethodsinitiative/4cat/master/.env
+
+# the two settings worth reading before you start it:
+# SERVER_BIND_ADDRESS=127.0.0.1 localhost only (the shipped default)
+# PUBLIC_PORT=80 the port the web UI lands on
+grep -E 'SERVER_BIND_ADDRESS|PUBLIC_PORT|DOCKER_TAG' .env
+
+docker compose up -d
+docker compose logs -f backend # first start builds indexes; wait for it
+
+# the one-time admin link is printed in the logs, not emailed
+docker compose logs frontend | grep -i 'create a new user\|token'
+```
+
+Creating a dataset, once it is up:
+
+```text
+1. http://localhost:80 -> "Create dataset".
+2. Pick the data source. Natively it collects 4chan, 8kun, Bluesky, Telegram, Tumblr
+ and TikTok (from a list of URLs). Everything else arrives as an upload.
+3. Set the query and the date range. Date range is the field people leave open and
+ then wonder why the job runs for a day -- bound it.
+4. Submit, then leave it. The backend queues and processes; the dataset appears in
+ your list with a status. Big Telegram or 4chan pulls take hours.
+5. On the finished dataset, run processors rather than exporting immediately:
+ word frequencies, co-word networks, time series, top posters, image downloads
+ and clustering. Each produces a child dataset you can chain further processors on.
+6. Export the one you want as CSV or NDJSON, and keep the parent dataset. The parent
+ is the thing you cannot re-create.
+```
+
+The platform coverage is the live constraint and it changes: X/Twitter, Instagram, LinkedIn,
+Threads and Pinterest are no longer collected by 4CAT directly — they come in through Zeeschuimer
+below, or as an upload from elsewhere. Datasets are big, and the Docker volumes will fill a disk
+without warning, so check free space before a large pull. And a 4CAT capture is a snapshot of what
+the platform served at that moment: edits and deletions after the capture are invisible to it,
+which is a feature when you want the original and a trap when you assume it reflects the live site.
+
+### Zeeschuimer
+
+The companion extension, and the current answer to "how do I get X, Instagram or TikTok data into
+4CAT". It watches the data the platform sends to your browser as you scroll, and keeps it. No API,
+no scraping requests of its own — it records what your normal browsing already fetched.
+
+```text
+1. Firefox only. Signed .xpi builds are on the project's releases page; Chrome is
+ not supported.
+2. Install, then open the extension and enable the platform you are about to browse.
+ Supported: TikTok, Instagram, Threads, X/Twitter, Pinterest, Gab, Truth Social,
+ 9gag, Imgur, Douyin and RedNote/Xiaohongshu.
+3. Browse normally -- a profile, a hashtag, a search. The item counter in the
+ extension ticks up as content loads. Content you never scrolled past is not
+ captured, because it was never sent to your browser.
+4. Export as NDJSON or CSV, or put your 4CAT URL in the extension and upload straight
+ into it as a dataset.
+5. Record how you browsed: which profile, which tab, how far you scrolled. That is
+ the sampling frame, and without it the dataset's coverage is undocumented.
```
-Self-hosting matters here: your dataset stays yours, and you can re-run analysis without
-re-collecting.
+This is manual collection with automatic recording, so it does not scale and it does not run
+unattended — which is also why it survives platform changes better than scrapers do. It captures
+through your logged-in session, so everything you collect is attributable to that account: use a
+research profile in a separate Firefox container or profile, never your own. Platform support needs
+constant maintenance and individual platforms do break; check the releases page before assuming the
+extension is at fault rather than the site.
+
+### changedetection.io
+
+Self-hosted page watching. This is the one to run when the thing you need to notice is an edit — a
+bio rewritten, a post deleted, an officer removed from a company page, a policy document quietly
+amended, a price or a staff list changing.
+
+```bash
+# docker, bound to localhost
+docker run -d --restart always -p "127.0.0.1:5000:5000" \
+ -v datastore-volume:/datastore \
+ --name changedetection.io dgtlmoon/changedetection.io
+
+# or as a package, if you would rather not run a container
+pip3 install changedetection.io
+changedetection.io -d /path/to/empty/data/dir -p 5000
+```
+
+```text
+1. http://localhost:5000 -> "Add new watch", paste the URL.
+2. Open "Edit" and set the filter before you set anything else. CSS selector, XPath,
+ JSONPath or jq -- point it at the narrowest element carrying the signal. Watching
+ a whole page means alerting on rotating ads, view counters and timestamps.
+ The "Visual Selector" tab picks the element by clicking it.
+3. Set "Ignore text" for the lines you know churn: relative dates, "N views",
+ cookie-banner text.
+4. Recheck time: per-watch. Daily for most things. The default applies to every watch
+ you add, so set it low once rather than per watch.
+5. Notifications: Discord, email, Slack, Telegram or a webhook via Apprise. A webhook
+ into your own notes or ticketing is the one that leaves a record.
+6. For a page that renders client-side, enable the Playwright/Sockpuppetbrowser
+ fetcher for that watch -- the default plain fetch sees an empty shell and will
+ report "no change" forever.
+7. "Browser Steps" handles a page behind a form: click, fill, submit, then diff what
+ comes back.
+8. Every change is stored as a snapshot with a diff view, which is the part that makes
+ it evidence rather than an alert.
+```
+
+Running it yourself means your watch list is not a third party's business record, which matters
+when the watch list itself is sensitive. The cost is that a JS-heavy watch runs a real browser and
+is far heavier than a text diff — a dozen of those on a small VPS will struggle. And a watch only
+sees what an unauthenticated fetch from your server sees: a page that needs a login, or that
+geo-varies, needs Browser Steps or will silently watch the wrong thing.
+
+### Distill
+
+The hosted equivalent, for when you will not run infrastructure. Same idea, less setup, and a free
+tier that is genuinely usable for a handful of watches: **25 monitors total but only 5 in the
+cloud, a 6-hour minimum cloud interval, 1,000 cloud checks a month, 2 devices, and 30 email alerts
+a month**. Local monitors in the browser extension are unlimited but only run while the browser is
+open.
+
+```text
+1. Browser extension or distill.io. On the page you want, click the extension and
+ select the region -- it generates the selector for you.
+2. Choose local (runs in your browser, unlimited, only while open) or cloud (runs
+ without you, capped as above). For anything that matters, cloud.
+3. Set the check interval. Anything under 6 hours is a paid feature on the free tier,
+ and six-hourly is adequate for almost all of this work anyway.
+4. Set the condition, not just "any change" -- Distill supports text conditions, so
+ "alert when the number changes" beats "alert when the page differs".
+5. Alerts to email or webhook; the email allowance is the binding constraint on free.
+6. Export the watch list as JSON periodically. It is the only part that is painful to
+ rebuild.
+```
+
+Your watch list and every snapshot live on their servers, which is the trade for not running
+anything. For a target that could plausibly subpoena or compromise a third party, that is the wrong
+trade and `changedetection.io` is the answer instead. The free tier's 6-hour floor also means you
+will miss a post that goes up and comes down inside a window, which is precisely the kind of
+deletion worth catching — if that is the scenario, self-host and poll faster.
+
+### yt-dlp with a download archive
+
+For recurring capture of a channel's output, the archive file is the whole trick: it records the ID
+of everything already fetched, so the next run picks up only what is new. That turns a one-off
+download into a monitor you can cron.
+
+```bash
+pipx install yt-dlp
+
+# first run: establish the archive. --break-on-existing stops as soon as it meets
+# something already recorded, so later runs walk only the new items
+yt-dlp --download-archive archive.txt --break-on-existing --lazy-playlist \
+ -o '%(upload_date)s-%(id)s.%(ext)s' \
+ 'https://youtube.com/@channel/videos'
+
+# metadata-only monitoring: no video files, just the record that it existed
+yt-dlp --download-archive seen.txt --break-on-existing \
+ --skip-download --write-info-json --write-thumbnail \
+ 'https://youtube.com/@channel/videos'
+
+# bound it by date instead, for a backfill of a known window
+yt-dlp --dateafter 20260101 --datebefore 20260401 --download-archive archive.txt URL
+
+# subtitles, which turn a channel's output into greppable text as it arrives
+yt-dlp --download-archive seen.txt --break-on-existing --skip-download \
+ --write-auto-subs --sub-langs en 'https://youtube.com/@channel/videos'
+
+# throttle it. these two flags are the difference between a monitor and a nuisance
+yt-dlp --sleep-requests 2 --sleep-interval 10 --max-sleep-interval 30 \
+ --download-archive archive.txt URL
+
+# several channels in one run, each stopping at its own first-seen item
+yt-dlp --break-per-input --break-on-existing --download-archive archive.txt \
+ -a channels.txt
+
+# a cap, so a misconfigured run cannot pull a thousand files overnight
+yt-dlp --max-downloads 50 --download-archive archive.txt URL
+```
-## Change detection
+The archive file is state: back it up, and never delete it to "start fresh" unless you mean to
+re-download everything. `--break-on-existing` assumes the listing is newest-first, which it is for
+channels and playlists and is not for some search result pages — on those, drop it and let the
+archive do the skipping. Extractors break when platforms change their pages, so `pipx upgrade
+yt-dlp` before blaming a URL, and a cron job that has silently failed for three weeks is worse than
+no monitoring at all, so alert on non-zero exits. Adding `--cookies-from-browser` attaches a real
+session to every request and makes the whole monitor attributable; avoid it unless the content
+genuinely requires a login. Full metadata usage is on
+[Social Media Platforms](/sheets/osint/social-media-platforms).
-[Distill](https://distill.io/) watches a page or a page region and alerts on change. Good for
-noticing when a target edits a bio, deletes a post, changes a company officer list, or quietly
-updates a policy document.
+### Native feeds, before you reach for a bridge
-Point it at the narrowest element that carries the signal, not the whole page, or you will drown in
-alerts from rotating ads and timestamps.
+Several platforms still publish perfectly good feeds that nobody uses because everyone assumes RSS
+died. These are keyless, stable, cheap to poll and attach to no account — the best monitoring
+surface available, where it exists.
+
+```bash
+# YouTube channel, by channel ID (the UC... form; @handles do not work here)
+curl -s 'https://www.youtube.com/feeds/videos.xml?channel_id=UCXuqSBlHAE6Xw-yeJA0Tunw'
+
+# a subreddit's new posts
+curl -s -A 'research-monitor/1.0 (contact: you@example.org)' \
+ 'https://www.reddit.com/r/osint/new/.rss'
+
+# one Reddit user's activity
+curl -s -A 'research-monitor/1.0 (contact: you@example.org)' \
+ 'https://www.reddit.com/user/someuser/.rss'
+
+# any Mastodon account, on any instance
+curl -s 'https://mastodon.social/@Gargron.rss'
+
+# a Bluesky profile
+curl -s 'https://bsky.app/profile/bsky.app/rss'
+
+# a GitHub user's public activity, which dates account behaviour precisely
+curl -s 'https://github.com/bellingcat.atom'
+
+# pull just the timestamps, to see cadence without reading content
+curl -s 'https://www.youtube.com/feeds/videos.xml?channel_id=UCXuqSBlHAE6Xw-yeJA0Tunw' \
+ | grep -oE '<published>[^<]+' | sed 's/<published>//'
+```
+
+Reddit rate-limits these hard and will return HTTP 429 to an anonymous, default-User-Agent client
+within a handful of requests; a descriptive User-Agent and a gap of several seconds between calls
+fixes it, and Reddit's `search.rss` endpoint is throttled more aggressively than the subreddit and
+user feeds. YouTube's feed carries only the most recent entries, so it monitors but does not
+backfill. Bluesky's profile RSS covers posts and not replies or likes — the AT Protocol endpoints
+on [Social Media Platforms](/sheets/osint/social-media-platforms) go deeper. And X/Twitter publishes
+nothing of this kind: `snscrape` has been non-functional against it since the 2023 access changes
+and the public Nitter instances are largely dead, so there is no quiet monitoring surface for X at
+all — only a logged-in session, with everything that implies.
+
+### RSSHub and RSS-Bridge
+
+For the platforms that killed their feeds. Both are self-hosted services that scrape a site and
+re-emit it as RSS or Atom, so your monitor still speaks one protocol no matter how many platforms
+are behind it.
+
+```bash
+# RSSHub -- the larger of the two, ~1000 routes
+docker run -d --name rsshub -p 1200:1200 diygod/rsshub
+# or with the compose file, which adds redis caching and a browser for JS sites
+wget https://raw.githubusercontent.com/DIYgod/RSSHub/master/docker-compose.yml
+docker compose up -d
+
+# routes are paths. a Telegram public channel:
+curl -s 'http://localhost:1200/telegram/channel/awesomeRSSHub'
+# the route catalogue lives at docs.rsshub.app/routes/
+
+# RSS-Bridge -- fewer routes (~450) but simpler, and a web UI that builds the URL
+docker create --name=rss-bridge --publish 3000:80 \
+ --volume $(pwd)/config:/config rssbridge/rss-bridge
+docker start rss-bridge
+
+# bridges are query parameters rather than paths
+curl -s 'http://localhost:3000/?action=display&bridge=<BridgeName>&format=Atom'
+```
+
+Self-host both. The public `rsshub.app` demo instance returns HTTP 403 to ordinary requests and is
+not a reliable backend for anything you depend on, and a shared public instance makes your watch
+list someone else's log file either way. Expect individual routes to break: they are scrapers
+wearing an RSS hat, and when a platform changes its markup the route returns an empty feed rather
+than an error — which reads exactly like "the target stopped posting". Check periodically that a
+route still returns items, and never conclude silence from an empty bridge feed without confirming
+against the site.
+
+### curl and jq on a schedule
+
+For a public JSON endpoint, a few lines beat any framework. The pattern that matters is the
+watermark: store the timestamp of the newest item you have seen, and ask only for things after it.
+
+```bash
+# poll Arctic Shift for new Reddit posts in a subreddit since the last run.
+# full Arctic Shift usage is on the social media platforms sheet; this is the
+# monitoring shape of it
+AS='https://arctic-shift.photon-reddit.com/api'
+STATE=~/monitor/last_seen_osint
+SINCE=$(cat "$STATE" 2>/dev/null || echo 0)
+
+curl -sS --fail -G "$AS/posts/search" \
+ --data-urlencode 'subreddit=osint' \
+ --data-urlencode "after=$SINCE" \
+ --data-urlencode 'sort=asc' \
+ --data-urlencode 'limit=100' \
+ -o /tmp/new.json || { echo "poll failed $(date -u +%FT%TZ)" >> ~/monitor/errors.log; exit 1; }
+
+# advance the watermark only on a successful fetch, or a failure silently
+# becomes a gap you never notice
+jq -r '.data[-1].created_utc // empty' /tmp/new.json | grep . && \
+ jq -r '.data[-1].created_utc' /tmp/new.json > "$STATE"
+
+# append raw, then derive. the raw file is the thing you cannot regenerate
+cat /tmp/new.json >> ~/monitor/osint-raw.ndjson
+jq -r '.data[] | [.created_utc, .author, .title] | @tsv' /tmp/new.json
+```
+
+```bash
+# a Bluesky author feed, keyless, as a second example of the same shape
+curl -sS --fail 'https://public.api.bsky.app/xrpc/app.bsky.feed.getAuthorFeed?actor=bsky.app&limit=50' \
+ | jq -r '.feed[] | [.post.indexedAt, .post.record.text] | @tsv'
+
+# crontab: daily, at a minute that is not :00, with failures mailed to you
+# 37 6 * * * /home/you/monitor/poll-osint.sh >> /home/you/monitor/run.log 2>&1
+```
+
+`--fail` is the flag that turns a silent HTML error page into a non-zero exit, and without it your
+monitor will happily append a Cloudflare block page to its dataset for a month. Advance the
+watermark only after a confirmed good response, log every failure with a timestamp, and treat the
+log as part of the dataset — the difference between "they went quiet" and "we were blocked" lives
+there and nowhere else. Note that Bluesky's `getAuthorFeed` is open without a token but
+`searchPosts` on the same public host returns 403 unauthenticated, which is the sort of asymmetry
+worth checking per endpoint before you build a schedule around it.
+
+### Arctic Shift for Reddit history
+
+The live successor to Pushshift, and the only practical way to get Reddit content that has since
+been deleted or edited. For monitoring it plays a specific role: it is the backfill that gives your
+forward-looking poll a baseline, and the thing you check when a post you captured disappears.
+
+```bash
+AS='https://arctic-shift.photon-reddit.com/api'
+
+# the baseline: what this account did before you started watching it
+curl -s -G "$AS/comments/search" --data-urlencode 'author=some_user' \
+ --data-urlencode 'sort=asc' --data-urlencode 'limit=100' \
+ | jq -r '.data[] | [.created_utc, .subreddit] | @tsv'
+
+# posting cadence by hour of day, which is what exposes a coordinated account
+curl -s -G "$AS/posts/search" --data-urlencode 'author=some_user' \
+ --data-urlencode 'limit=500' --data-urlencode 'fields=created_utc' \
+ | jq -r '.data[].created_utc' \
+ | python3 -c 'import sys,datetime,collections; c=collections.Counter(datetime.datetime.fromtimestamp(int(l), datetime.UTC).hour for l in sys.stdin); [print(f"{h:02d} {c[h]}") for h in range(24)]'
+
+# did a post you captured actually get removed, or did you lose it
+curl -s -G "$AS/posts/ids" --data-urlencode 'ids=t3_abc123' | jq '.data[0] | {title, selftext, removed_by_category}'
+```
+
+Keyword search needs an accompanying `author`, `subreddit`, `link_id` or `parent_id` — bare keyword
+sweeps are refused, and so are very broad author or subreddit queries. The archive holds what it
+ingested at the time, so an edit made after ingestion is invisible and a post deleted within
+seconds may never have been captured at all; a record here proves the text was published, not that
+it is live. Bulk dumps for offline work are published separately from the API. Query syntax in full
+is on [Social Media Platforms](/sheets/osint/social-media-platforms).
## Narrative and coordination analysis
- **Posting-time clustering** exposes networks. Accounts that post within seconds of each other,
- repeatedly, are coordinated.
+ repeatedly, are coordinated. The `jq` plus histogram pattern above is enough to see it.
- **Identical phrasing across accounts** is the strongest single indicator of copy-paste campaigns.
- **Follower overlap** between accounts is more telling than follower count.
-- [The Information Laundromat](https://information-laundromat.com/) compares content across sites to
- find where the same text is being syndicated.
+- [The Information Laundromat](https://informationlaundromat.com/) compares content and technical
+ metadata — ad IDs, analytics tags, registration details, code structure — across sites, to find
+ where the same text is syndicated and which sites share infrastructure. Built by the Alliance for
+ Securing Democracy, which merged into the Institute for Strategic Dialogue on 1 January 2026; the
+ tool remains live at the address above. The older hyphenated domain no longer resolves.
+- Graphing and clustering what you have collected is on
+ [Data Analysis & Visualisation](/sheets/osint/data-analysis-and-visualisation).
## Tool reference
@@ -72,17 +474,73 @@ alerts from rotating ads and timestamps.
| [Maltego Graph](https://www.maltego.com/downloads/) | Maltego Graph is an investigation platform that combines two things at once: (1) It acts as a search tool, and (2) It creates a graph establishing links… | partly free |
| [Pinpoint](https://journaliststudio.google.com/pinpoint/about) | A tool by Google to catalogue uploaded documents and files, providing automated text recogntion, indexing, audiotranscriptions and other (AI-powered)… | free |
| [Time.Graphics](https://time.graphics) | A tool for creating, visualizing, and managing timelines online. | partly free |
+| [Zeeschuimer](https://github.com/digitalmethodsinitiative/zeeschuimer) | Firefox extension that records social media data as you browse it, for import into 4CAT. | free |
+| [changedetection.io](https://changedetection.io/) | Self-hosted page-change monitoring with CSS/XPath/jq filters and diff history. | free |
+| [Distill](https://distill.io/) | Hosted page-change monitoring, browser extension plus cloud checks. | partly free |
## Pitfalls
+- **Dead tools still appear in tutorials.** `snscrape` and the public Nitter instances do not work
+ against X. Bing's Visual Search API and the rest of the Bing Search family were decommissioned in
+ August 2025. Check a tool is alive before you design a collection around it.
- **Rate limits end collections mid-run.** Check for gaps before you analyse; a missing day looks
like silence rather than a failure.
+- **An empty feed is ambiguous.** A broken RSSHub route, a changed selector and a target who
+ stopped posting all look identical downstream. Monitor your monitors.
- **Sampling bias reads as a finding.** If a tool only reaches accounts above some follower count,
- its "network" is an artefact of that cutoff.
+ or only what you happened to scroll past, its "network" is an artefact of that cutoff.
- **Coordination needs a baseline.** Fans of the same thing post about it at the same time. Compare
against normal behaviour for the topic before calling it inauthentic.
+- **Monitoring is contact.** Scheduled requests are a pattern; a logged-in monitor is an
+ attributable pattern. Decide what that costs before you start, not after.
- **Storage and legality.** Bulk personal data attracts data-protection obligations even when every
- individual item was public.
+ individual item was public, and a dataset is harder to delete than to collect.
+
+## Worked example
+
+One datum: a single Telegram channel name, from a screenshot, alleged to be seeding a story that
+later appeared on a cluster of news-like websites.
+
+1. **Baseline before watching.** Pull the channel's existing history into 4CAT as a Telegram
+ dataset with a bounded date range. You now know its normal posting cadence and vocabulary, which
+ is what any later claim of a surge has to be measured against.
+2. **Find the forward-looking surface.** Telegram has no public feed, so an RSSHub route
+ (`/telegram/channel/<name>` on your own instance) gives you a pollable endpoint. Verify it
+ returns items today, so that an empty feed next month means something.
+3. **Schedule it honestly.** A daily `curl` with a stored watermark, a descriptive User-Agent, and
+ a failure log. Not hourly — a story that takes days to syndicate does not need fifteen-minute
+ resolution, and the slower poll survives longer.
+4. **Watch the downstream sites for edits, not just posts.** Each suspected site gets a
+ `changedetection.io` watch filtered to the article body, so a quietly amended paragraph or a
+ removed byline raises an alert with a stored diff. This is the part that a collection-only
+ approach misses entirely.
+5. **Capture the video claims properly.** The channel posts clips; `yt-dlp --download-archive`
+ with `--write-info-json` on its linked channels picks up only what is new each day and records
+ `upload_date` for each, giving every clip a latest-possible date.
+6. **Test the syndication claim.** Feed one article URL to the Information Laundromat. Content
+ similarity tells you the text is shared; the technical indicators — a common analytics ID across
+ four of the sites — tell you something stronger, because wording can be copied by anyone and a
+ shared tracking ID usually cannot.
+7. **Cross-check the Reddit leg.** Arctic Shift for posts linking those domains, grouped by
+ subreddit and by hour, shows whether the amplification is a handful of accounts on a schedule or
+ genuine spread.
+8. **Archive as you go.** Every page you will cite goes to a snapshot at the time you saw it, per
+ [Archiving & Evidence](/sheets/osint/archiving-and-evidence) — these sites edit and disappear,
+ which is the behaviour you are documenting.
+
+What you can assert: a dated posting history for the channel, dated first appearances on each
+downstream site, stored diffs of any subsequent edits, and a shared technical indicator linking
+some of those sites.
+
+What would falsify it: an RSSHub route that broke silently mid-period, which would turn a real gap
+into an apparent one — so the run log, showing a successful fetch every day, is doing as much
+evidential work as the data. A shared analytics ID that turns out to belong to a common CMS
+template, or to an agency that serves unrelated clients, would break the infrastructure link; check
+what else carries that ID before leaning on it.
+
+## Broader catalogues
+
+- [Social Media OSINT](https://tools.osintnewsletter.com/tool-categories/social-media-osint)
## Sources
diff --git a/src/pages/credits.astro b/src/pages/credits.astro
@@ -134,9 +134,8 @@ SOFTWARE.`;
<p>
<a href="https://ired.team" target="_blank" rel="noopener">ired.team</a> — the red teaming notes of
<strong>Mantvydas Baranauskas</strong> (<a href={iredSource.authorUrl} target="_blank" rel="noopener">@mantvydasb</a>) —
- is the reference behind a number of sheets in Exploitation, Password Attacks, Privilege Escalation,
- Tunneling & Pivoting and DFIR. His write-ups on process injection, defense evasion and persistence are
- some of the most careful hands-on documentation of those techniques anywhere.
+ is the reference behind sheets in Exploitation and DFIR. His write-ups on process injection and on
+ Windows internals are some of the most careful hands-on documentation of those subjects anywhere.
</p>
<p>
<strong>ired.team publishes no licence.</strong> That grants no right to copy it, so nothing from it is