archiving-and-evidence.md (40425B)
1 --- 2 title: "Archiving & Evidence Preservation" 3 description: "Capture sources so they survive deletion, with the hashes and timestamps that make a capture defensible." 4 category: osint 5 subcategory: "Archiving & Analysis" 6 tags: [osint, archiving, evidence, preservation] 7 tools: [wget, browsertrix-crawler, yt-dlp, auto-archiver, wayback, archive-today, opentimestamps, openssl, exiftool, shasum] 8 difficulty: intermediate 9 updated: 2026-10-04 10 references: 11 - name: "Bellingcat's Online Investigation Toolkit" 12 url: "https://bellingcat.gitbook.io/toolkit" 13 author: "Bellingcat" 14 license: none 15 relation: derived 16 note: "Tool catalogue: names, descriptions, cost flags and links for this area." 17 - name: "OSINT Newsletter Tools Library" 18 url: "https://tools.osintnewsletter.com" 19 author: "The OSINT Newsletter" 20 license: none 21 relation: derived 22 note: "Second tool catalogue, cross-checked against the above." 23 --- 24 25 ## What this covers 26 27 Capturing a source so that it still exists after someone deletes it, and so that you can show the 28 copy you hold is the copy you took. The capture itself is the easy half; the hash, the third-party 29 timestamp and the written provenance are what make it worth anything to an editor, a lawyer or a 30 court. Everything below is organised around one question: what does this step actually prove? 31 32 ## Method 33 34 1. **Check what already exists before you capture.** A Wayback snapshot from before the content 35 became interesting is worth more than one you took after. Query the availability and CDX APIs 36 first; if there is an older capture, that is your baseline and the diff against today is itself a 37 finding. 38 2. **Capture to something you do not control, and to something you do.** A third-party archive 39 supplies a timestamp nobody can accuse you of forging. A local copy survives the archive being 40 rate-limited, robots-excluded or taken down. Neither substitutes for the other. 41 3. **Pick the capture tool by what the page is.** Static HTML and media: `wget --warc-file`. 42 JavaScript-rendered, infinite-scroll or login-walled: a real browser driving the capture, which 43 in practice means `browsertrix-crawler` producing WACZ. Platform video: `yt-dlp`, which gets you 44 the metadata sidecar a screen recording never will. 45 4. **Hash on acquisition, before you look at the file.** A hash taken after you cropped, converted 46 or re-saved proves something about your working copy and nothing about the original. 47 5. **Timestamp the manifest, not every file.** One hash over the list of hashes, submitted to 48 OpenTimestamps or an RFC 3161 authority, binds the whole collection to a point in time at one 49 cost. That is the step that converts "I have a hash" into "I had this hash on that date". 50 6. **Write the provenance down in the same commit as the files.** URL, UTC time of capture, the 51 tool and version, who handed it to you, and how you reached the page. The route to a finding is 52 part of the finding, and it is the part you will forget first. 53 7. **Get a second independent capture.** Two archives of the same URL taken by different 54 infrastructure, or the same video from a second uploader, is the difference between a claim and 55 a corroborated claim. Two copies that both trace to one upload are one copy. 56 57 Judgement calls worth naming: submitting a URL to a public archive creates a public record that 58 somebody is interested in it, so on a target that watches its referrers and its Wayback entries, 59 capture locally first and submit later. And a capture of a page you reached while logged in 60 contains your session — strip or redact before you share the WARC. 61 62 ## Key tools 63 64 ### wget (WARC capture) 65 66 The standard way to get a byte-level record of an HTTP exchange rather than a rendered picture of 67 it. A WARC stores the request and response headers alongside the body for every resource fetched, 68 with a SHA-1 digest per record, so the file is a transcript of the conversation and not just its 69 outcome. Use it whenever the page is server-rendered; it is the most portable evidence format in 70 this area and every replay tool reads it. 71 72 ```bash 73 brew install wget # or: apt install wget 74 ``` 75 76 ```bash 77 # the baseline capture: WARC transcript plus a CDX index of what went into it 78 wget --warc-file=capture-001 --warc-cdx \ 79 --page-requisites --adjust-extension --no-parent \ 80 'https://example.com/page' 81 82 # stamp the operator and case into the warcinfo record, so the file self-documents 83 wget --warc-file=capture-002 --warc-cdx \ 84 --warc-header="operator: J. Investigator" \ 85 --warc-header="description: case-2026-014, post cited in filing" \ 86 --warc-header="robots: off" \ 87 --page-requisites --adjust-extension 'https://example.com/page' 88 89 # leave it uncompressed while you are inspecting it; grep works on a plain WARC 90 wget --warc-file=capture-003 --no-warc-compression 'https://example.com/page' 91 92 # pull in the CDN that actually serves the images, or the capture renders blank 93 wget --warc-file=capture-004 --warc-cdx \ 94 --page-requisites --convert-links --adjust-extension \ 95 --span-hosts --domains example.com,cdn.example.com \ 96 'https://example.com/page' 97 98 # a whole section, politely: the delay is what stops you being blocked mid-capture 99 wget --warc-file=capture-005 --warc-cdx --recursive --level=2 --no-parent \ 100 --wait=2 --random-wait --limit-rate=500k \ 101 -U 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0 Safari/537.36' \ 102 'https://example.com/newsroom/' 103 104 # a list of URLs in one WARC, split at 1GB so the files stay movable 105 wget --warc-file=capture-006 --warc-max-size=1G --warc-cdx -i urls.txt 106 107 # second pass over the same site without re-storing what you already have 108 wget --warc-file=capture-007 --warc-dedup=capture-005.cdx --recursive --no-parent \ 109 'https://example.com/newsroom/' 110 111 # read back what the capture contains, without a replay tool 112 zcat capture-001.warc.gz | grep -a '^WARC-Target-URI:' | sort -u 113 awk 'NR>1 {print $5, $6, $1}' capture-001.cdx # status, SHA-1 digest, URL 114 ``` 115 116 Read the CDX file: the digest column is a base32 SHA-1 of the response payload, which gives you a 117 free per-resource integrity check, and the status column tells you which requisites 404'd. A 118 capture whose CDX is full of `302` and `403` rows did not get the page. 119 120 What this does not prove: the WARC records what *your* client received from *an* IP at that moment. 121 It does not prove the page looked that way to anyone else, does not survive geo-targeted or 122 personalised content, and carries no third-party attestation of the time — the `WARC-Date` is your 123 own clock. `--convert-links` rewrites the mirrored files on disk, not the WARC records, so the 124 transcript stays pristine while the browsable copy is usable; keep both. Above all, `wget` executes 125 no JavaScript, so on a modern single-page application it faithfully archives an empty shell. If the 126 page needs a browser, use one. 127 128 ### browsertrix-crawler 129 130 A headless Chrome driven by Webrecorder's crawler, writing WARC and WACZ. This is the live, 131 maintained answer for anything JavaScript-rendered, where `wget` returns a loading spinner. It 132 scrolls, autoplays, runs site-specific behaviours for the big platforms, and extracts page text for 133 search. 134 135 The older Python crawlers in this niche have aged out: `wpull` carries a 2013–2016 copyright and no 136 release since 2019, and `grab-site` still pins Python 3.7/3.8 and was dropped from nixpkgs after 137 23.05. Neither is a reasonable thing to stand behind in 2026. Treat them as read-only history and 138 use the crawler below. 139 140 ```bash 141 docker pull webrecorder/browsertrix-crawler 142 ``` 143 144 ```bash 145 # single page, as WACZ, with text extracted for full-text search 146 docker run -v "$PWD/crawls:/crawls/" -it webrecorder/browsertrix-crawler crawl \ 147 --url 'https://example.com/post/123' --scopeType page \ 148 --generateWACZ --text to-pages --collection case-2026-014 149 150 # the page plus whatever it links to on the same host, two levels deep 151 docker run -v "$PWD/crawls:/crawls/" -it webrecorder/browsertrix-crawler crawl \ 152 --url 'https://example.com/newsroom/' --scopeType host --depth 2 \ 153 --generateWACZ --collection newsroom 154 155 # a feed that only loads on scroll: give the behaviours time to finish 156 docker run -v "$PWD/crawls:/crawls/" -it webrecorder/browsertrix-crawler crawl \ 157 --url 'https://example.com/feed' --scopeType page-spa \ 158 --behaviors autoscroll,autoplay,autofetch,siteSpecific \ 159 --behaviorTimeout 300 --generateWACZ --collection feed 160 161 # several URLs from a file, one collection, so the manifest covers the set 162 docker run -v "$PWD/crawls:/crawls/" -it webrecorder/browsertrix-crawler crawl \ 163 --urlFile /crawls/urls.txt --scopeType page --generateWACZ --collection batch-01 164 165 # the WACZ is a zip: look inside before you trust it 166 unzip -l crawls/collections/case-2026-014/case-2026-014.wacz 167 unzip -p crawls/collections/case-2026-014/case-2026-014.wacz datapackage-digest.json | jq . 168 169 # the crawl log is JSON Lines — this is where blocked and failed pages surface 170 jq -r 'select(.logLevel=="error") | [.timestamp,.message] | @tsv' \ 171 crawls/collections/case-2026-014/logs/*.log 172 ``` 173 174 A WACZ is a WARC plus an index, a page list and a signed digest manifest, and it opens in 175 [ReplayWeb.page](https://replayweb.page/) with no server. That makes it the format to hand to 176 someone who is not going to install anything. Check `datapackage-digest.json` after every crawl: 177 if the crawler was blocked, you get a valid WACZ containing a block page, and nothing about the 178 file itself says so. 179 180 What it does not tell you: a browser capture is a recording of one rendering session, which means 181 it includes whatever was personalised, A/B-tested or geo-fenced for that session. Running the same 182 crawl from a different exit and getting different content is a finding about the site, not an error. 183 Logged-in crawls bake the session into the archive — handle accordingly. 184 185 ### yt-dlp 186 187 Platform video and audio with its metadata intact. The point is not the video file; a screen 188 recording gets you that. The point is the `.info.json` sidecar, which carries uploader ID, upload 189 date, duration, view and comment counts, available formats and often the original title and 190 description before anybody edited them. 191 192 ```bash 193 brew install yt-dlp # or: pipx install yt-dlp 194 ``` 195 196 ```bash 197 # the provenance capture: video, metadata, description, thumbnail, subtitles 198 yt-dlp --write-info-json --write-description --write-thumbnail \ 199 --write-subs --sub-langs 'all' --no-mtime \ 200 -o '%(upload_date)s_%(uploader_id)s_%(id)s.%(ext)s' URL 201 202 # what formats exist, before you commit to one 203 yt-dlp -F URL 204 205 # the best single pre-merged file, so you archive bytes the platform served 206 yt-dlp -f 'best[ext=mp4]/best' --no-mtime URL 207 208 # metadata only — enough to date and attribute a clip without downloading it 209 yt-dlp --skip-download --write-info-json --write-thumbnail URL 210 211 # comments too; they routinely date an event more precisely than the upload does 212 yt-dlp --write-info-json --write-comments --skip-download URL 213 214 # a whole channel, newest first, with an archive file so re-runs are incremental 215 yt-dlp --write-info-json --no-mtime --download-archive seen.txt \ 216 -o '%(upload_date)s_%(id)s.%(ext)s' 'https://example.com/@account/videos' 217 218 # the intermediate pages the extractor fetched — useful when an extractor misreads a page 219 yt-dlp --write-pages --skip-download URL 220 221 # the fields you will actually quote, straight out of the sidecar 222 jq -r '[.id,.uploader_id,.upload_date,.timestamp,.duration,.view_count,.title] | @tsv' *.info.json 223 ``` 224 225 `--no-mtime` matters: without it the file's mtime is set from the upload date, which silently 226 overwrites your own acquisition timeline with the platform's claim. `timestamp` in the sidecar is 227 Unix epoch seconds and is usually more precise than `upload_date`, which is date-only and rendered 228 in the platform's timezone of choice. 229 230 What the sidecar does not establish: every field in it is the platform's assertion, repeated. 231 `upload_date` is when it was uploaded, never when it was filmed, and a re-upload resets it. 232 Extractors break when platforms change, so a field that is `null` may mean absent or may mean the 233 extractor lost it — check against the page. Rate limits and age or region gates will silently 234 truncate a channel pull; compare the count you got against the count the channel claims. 235 236 ### Wayback Machine, from the command line 237 238 Three separate endpoints, used at three different moments. Availability answers "is there already a 239 capture"; CDX answers "what captures exist, and did the content change between them"; Save Page Now 240 creates a new one. 241 242 ```bash 243 # is there a capture at all, and what is the nearest one to a date 244 curl -s 'https://archive.org/wayback/available?url=example.com/page×tamp=20240101' | jq . 245 246 # every capture of a URL, with the payload digest that reveals real edits 247 curl -s 'https://web.archive.org/cdx/search/cdx?url=example.com/page&output=json&fl=timestamp,original,statuscode,digest,length' | jq -r '.[] | @tsv' 248 249 # collapse consecutive identical captures: what is left is the change history 250 curl -s 'https://web.archive.org/cdx/search/cdx?url=example.com/page&output=json&fl=timestamp,digest&collapse=digest' | jq -r '.[] | @tsv' 251 252 # everything ever captured under a host, which is how you find deleted pages 253 curl -s 'https://web.archive.org/cdx/search/cdx?url=example.com&matchType=domain&output=json&fl=timestamp,original,statuscode&limit=500' | jq -r '.[] | @tsv' 254 255 # only the captures in a window, for a page that mattered on one specific day 256 curl -s 'https://web.archive.org/cdx/search/cdx?url=example.com/page&from=20260301&to=20260331&output=json&fl=timestamp,digest' | jq -r '.[] | @tsv' 257 258 # submit a new capture: SPN2, authenticated with Internet Archive S3-style keys 259 curl -s -X POST -H 'Accept: application/json' \ 260 -H "Authorization: LOW $IA_ACCESS_KEY:$IA_SECRET_KEY" \ 261 -d 'url=https://example.com/page' -d 'capture_all=1' -d 'capture_screenshot=1' \ 262 'https://web.archive.org/save' | jq . 263 264 # poll the job until it reports success, and keep the returned timestamp 265 curl -s -H "Authorization: LOW $IA_ACCESS_KEY:$IA_SECRET_KEY" \ 266 "https://web.archive.org/save/status/$JOB_ID" | jq '{status, timestamp, original_url}' 267 268 # fetch the archived copy itself, and hash what you got 269 curl -sL 'https://web.archive.org/web/20260301120000id_/https://example.com/page' | shasum -a 256 270 ``` 271 272 Get the key pair from [archive.org/account/s3.php](https://archive.org/account/s3.php) and export 273 it as `IA_ACCESS_KEY` / `IA_SECRET_KEY`. **The unauthenticated `GET /save/<url>` form is no longer 274 usable for work** — an anonymous request to it returns HTTP 429 rather than a capture. Scripts that 275 still rely on it fail silently, because a 429 body looks like a page. 276 277 The `id_` infix in that last URL asks for the original bytes without Wayback's navigation banner 278 injected, which is what you want if you intend to hash or diff the archived copy. The `digest` 279 column in CDX output is the payload hash: two captures sharing a digest are byte-identical, so 280 collapsing on it turns a thousand snapshots into the handful of moments the page actually changed. 281 That is the single most useful thing the CDX API does. 282 283 What Wayback does not give you: permanence. A site can retroactively exclude itself, and captures 284 do disappear. It also honours robots directives at crawl time, so absence of a capture is not 285 evidence the page did not exist. And a Save Page Now capture is still just a capture — it attests 286 that the Internet Archive saw that content at that time, which is a strong third-party claim about 287 time and a weak one about authenticity. 288 289 ### archive.today 290 291 A separate archive with a separate failure mode, which is exactly why you use it alongside Wayback. 292 It renders pages that Wayback flattens, ignores robots directives, and has a long record of not 293 removing things. There is no API and the submission form sits behind bot protection, so this one is 294 a browser job — do not script it. 295 296 ```text 297 1. Open https://archive.today (archive.ph and archive.is are the same service; 298 if one mirror is blocked where you are, try another) 299 2. Paste the URL into the lower box, "My url is alive and I want to archive its content" 300 3. Solve the challenge if you get one. Do not automate this step — it is the step 301 that gets the service to block your address. 302 4. Wait for the capture. The result URL looks like https://archive.ph/AbC12 303 and the page header shows the capture time in UTC. 304 5. Click "screenshot" in the header to get the full-page render as a separate 305 artefact, and save it. The HTML capture and the screenshot fail differently. 306 6. Record the short URL AND the capture time. The short URL is your citation. 307 7. Before submitting, use the upper box to search for existing captures of the 308 same URL — the same reason you query Wayback's CDX API first. 309 ``` 310 311 Two things to know. The service is deliberately opaque about its infrastructure and funding, which 312 means you should not treat it as your only copy of anything; keep the local WARC. And it fetches 313 the page itself, from its own address, so a submission does not leak your IP to the target — but it 314 does create a public, searchable record that someone archived that URL at that minute. 315 316 ### Hashing and the manifest 317 318 Hashes are the whole evidentiary argument. A hash recorded at acquisition and published or 319 timestamped separately lets you show, later, that the file you are producing is the file you took. 320 The manifest pattern — one file listing every hash, then one hash of the manifest — scales that to a 321 collection without timestamping a thousand files. 322 323 ```bash 324 # macOS ships shasum; coreutils ships sha256sum. Same output format. 325 shasum -a 256 evidence.mp4 326 sha256sum evidence.mp4 327 ``` 328 329 ```bash 330 # hash the whole capture tree in a stable order, so the manifest is reproducible 331 find ./capture -type f -print0 | sort -z | xargs -0 shasum -a 256 > MANIFEST.sha256 332 333 # the manifest's own hash: this single line is what you timestamp and publish 334 shasum -a 256 MANIFEST.sha256 | tee MANIFEST.sha256.sha256 335 336 # verify the collection, any time later 337 shasum -a 256 -c MANIFEST.sha256 338 339 # show only what failed, which is what you actually want on a large set 340 shasum -a 256 -c MANIFEST.sha256 2>/dev/null | grep -v ': OK$' 341 342 # a provenance record that travels with the files, in the same directory 343 cat > PROVENANCE.txt <<'EOF' 344 case: 2026-014 345 url: https://example.com/post/123 346 captured_at: 2026-10-04T04:07:56Z # UTC, from `date -u +%FT%TZ` 347 captured_by: J. Investigator 348 tooling: wget 1.25.0 --warc-file; yt-dlp 2026.08.19 349 route: linked from https://example.com/newsroom/ on 2026-10-03 350 notes: page required no login; no personalisation observed 351 EOF 352 353 # fold the provenance into the manifest so the two cannot drift apart 354 shasum -a 256 PROVENANCE.txt >> MANIFEST.sha256 355 356 # find duplicates across two collections by hash, not by filename 357 sort MANIFEST.sha256 other/MANIFEST.sha256 | awk '{print $1}' | sort | uniq -d 358 ``` 359 360 Use SHA-256. MD5 and SHA-1 still appear in forensic tooling and are fine as identifiers, but a 361 collision-prone hash invites an argument you do not need to have. Note that `find | sort` matters: 362 an unsorted manifest changes order between runs and filesystems, which makes two manifests of the 363 same files look different. 364 365 What a hash proves and does not: it proves the bytes have not changed since the hash was taken. It 366 says nothing about when they were taken, nothing about where they came from, and nothing about 367 whether the content is true. On its own a hash in your own notes is worth very little, because you 368 could have written it at any time — which is what the next section fixes. 369 370 ### Timestamping: OpenTimestamps and RFC 3161 371 372 The step that turns a hash into evidence. A timestamp is a third party's signed assertion that a 373 given digest existed before a given moment. Without it, your hash only proves internal 374 consistency; with it, you can show you held that exact file on that date and could not have 375 produced it later. 376 377 ```bash 378 pipx install opentimestamps-client # the `ots` command 379 # openssl is already present for the RFC 3161 route 380 ``` 381 382 ```bash 383 # OpenTimestamps: anchors the digest in the Bitcoin blockchain, free, no account 384 ots stamp MANIFEST.sha256 385 386 # what the proof currently contains: pending calendar attestations, then a block 387 ots info MANIFEST.sha256.ots 388 389 # calendars need an hour or so to get into a block; upgrade the proof afterwards 390 ots upgrade MANIFEST.sha256.ots 391 392 # verify. Full verification wants a local Bitcoin node; without one you are 393 # trusting the calendar, which is weaker but still a third party 394 ots verify MANIFEST.sha256.ots 395 396 # verify a digest you were given, without holding the file 397 ots verify -d 355c357708e4840d12b0d4284d9fe1a911c79c013953874f03068b81127f3363 MANIFEST.sha256.ots 398 399 # RFC 3161 route: build a timestamp query over the file's SHA-256 400 openssl ts -query -data MANIFEST.sha256 -sha256 -cert -out manifest.tsq 401 402 # send it to a Time Stamping Authority and keep the signed reply 403 curl -s -H 'Content-Type: application/timestamp-query' \ 404 --data-binary @manifest.tsq https://freetsa.org/tsr -o manifest.tsr 405 406 # read the token: the TSA's time, its identity and the digest it signed over 407 openssl ts -reply -in manifest.tsr -text 408 409 # verify the token against the TSA's chain — this is the check a reviewer repeats 410 curl -sO https://freetsa.org/files/cacert.pem 411 curl -sO https://freetsa.org/files/tsa.crt 412 openssl ts -verify -data MANIFEST.sha256 -in manifest.tsr \ 413 -CAfile cacert.pem -untrusted tsa.crt 414 ``` 415 416 A successful verification prints `Verification: OK`, and that line is the thing you cite. Keep the 417 `.tsr`, the CA chain you verified against, and the exact file — the token is signed over the digest, 418 so a reviewer needs the file to recompute it. 419 420 Which route to use: OpenTimestamps costs nothing, needs no account, and anchors to a public 421 blockchain, so the proof outlives any single company — but it is coarse (block granularity, roughly 422 an hour) and full verification wants a Bitcoin node. An RFC 3161 TSA gives you a precise, signed, 423 PKI-rooted token that institutional reviewers already recognise, at the price of trusting that TSA 424 and its certificate remaining validatable. Do both; they are cheap and they fail differently. 425 426 What a timestamp does not prove: that the content is authentic, that the capture was complete, or 427 that you obtained it lawfully. It proves the file existed no later than that moment. That is a 428 narrow claim, and stating it narrowly is what makes it credible. 429 430 ### Bellingcat Auto Archiver 431 432 The pipeline to set up when you are collecting continuously rather than case by case. It reads URLs 433 from a feeder, runs a platform-specific extractor, then runs enrichers that hash, thumbnail, 434 timestamp and push to Wayback, and writes the results to a database you can read as a case index. 435 It is the only tool here that performs the hash-and-timestamp discipline on every item without you 436 remembering to. 437 438 ```bash 439 pipx install auto-archiver 440 auto-archiver --help # the flag list is generated from the installed modules 441 ``` 442 443 Configuration is a single YAML file; the default path is `secrets/orchestration.yaml`. Only modules 444 named under `steps` are loaded, so the file is both configuration and a declaration of the pipeline. 445 446 ```yaml 447 # orchestration.yaml 448 steps: 449 feeders: [cli_feeder] 450 extractors: [generic_extractor] 451 enrichers: [hash_enricher, meta_enricher, thumbnail_enricher, 452 timestamping_enricher, opentimestamps_enricher] 453 databases: [csv_db, console_db] 454 storages: [local_storage] 455 formatters: [html_formatter] 456 457 hash_enricher: 458 algorithm: SHA-256 # the only alternative is SHA3-512 459 460 timestamping_enricher: 461 tsa_urls: 462 - http://timestamp.identrust.com 463 - http://zeitstempel.dfn.de 464 allow_selfsigned: false # leaving this false is the whole point of the module 465 466 local_storage: 467 save_to: ./archived 468 path_generator: url # flat | url | random 469 filename_generator: static # names the file after its hash 470 471 csv_db: 472 csv_file: ./archived/results.csv 473 474 logging: 475 level: INFO 476 file: ./archived/archiver.log 477 ``` 478 479 ```bash 480 # one URL through the whole pipeline 481 auto-archiver --config orchestration.yaml 'https://example.com/post/123' 482 483 # several at once; the CSV gets one row per URL 484 auto-archiver --config orchestration.yaml 'https://example.com/a' 'https://example.com/b' 485 486 # no config file at all, for a one-off capture with sane defaults 487 auto-archiver --mode simple 'https://example.com/post/123' 488 489 # override the enricher list without editing the file 490 auto-archiver --config orchestration.yaml \ 491 --enrichers hash_enricher timestamping_enricher opentimestamps_enricher \ 492 'https://example.com/post/123' 493 494 # push to Wayback as part of the run, with the same Internet Archive key pair 495 auto-archiver --config orchestration.yaml \ 496 --enrichers hash_enricher wayback_extractor_enricher \ 497 --wayback_extractor_enricher.key "$IA_ACCESS_KEY" \ 498 --wayback_extractor_enricher.secret "$IA_SECRET_KEY" \ 499 'https://example.com/post/123' 500 501 # feed a column of URLs from a spreadsheet export 502 auto-archiver --config orchestration.yaml \ 503 --feeders csv_feeder --csv_feeder.files urls.csv --csv_feeder.column link 504 505 # turn up logging when an extractor silently returns nothing 506 auto-archiver --config orchestration.yaml --logging.level DEBUG 'https://example.com/post/123' 507 508 # the results CSV is the case index: one row per URL, with hashes and archive links 509 column -s, -t < ./archived/results.csv | less -S 510 ``` 511 512 Two enrichers do the evidentiary work and they are not interchangeable. `timestamping_enricher` 513 aggregates the item's file hashes into one text file and gets an RFC 3161 token over it from each 514 TSA in the list; `opentimestamps_enricher` anchors the same digest via the OpenTimestamps 515 calendars. Note that `freetsa.org` is deliberately absent from the shipped TSA defaults because its 516 certificate does not validate against system authorities — you can add it, but only together with 517 `allow_selfsigned: true`, which weakens exactly the property you wanted. 518 519 The `gsheet_feeder_db` module is what makes this a team tool: analysts paste URLs into a Google 520 Sheet, the archiver writes the hash, the stored path and the archive links back into adjacent 521 columns. The `wacz_extractor_enricher` shells out to Docker for browser-based capture, and 522 `wayback_extractor_enricher` needs the same key pair as Save Page Now. 523 524 Watch the version. The client checks on startup and says so when it is behind, and per-platform 525 extractors break often enough that an old release produces empty captures that look like 526 successes. After every run, scan the results CSV for rows that carry a hash but no media. 527 528 ### ExifTool, for capture-time metadata 529 530 Covered in depth on [Image & Video Forensics](/sheets/osint/image-video-forensics); the archiving 531 use is narrower. You are not asking whether the file is manipulated, you are recording what the 532 file asserted about itself at the moment it entered your custody, so that a later disagreement is 533 about the metadata rather than about your handling. 534 535 ```bash 536 # the full tag dump, archived next to the hash and never edited again 537 exiftool -j -g1 -a -u capture/evidence.mp4 > capture/evidence.mp4.exif.json 538 539 # the fields that date and attribute a capture, in one line 540 exiftool -CreateDate -ModifyDate -FileModifyDate -GPSPosition -Make -Model -Software \ 541 capture/evidence.mp4 542 543 # every timestamp in the file, grouped by where it came from 544 exiftool -time:all -a -G1 -s capture/evidence.mp4 545 546 # the whole capture tree as one CSV, which doubles as a collection inventory 547 exiftool -r -csv -FileName -FileSize -MIMEType -CreateDate -GPSPosition ./capture/ > inventory.csv 548 549 # flag files whose structure does not match their declared format 550 exiftool -r -validate -warning -a ./capture/ 551 552 # strip metadata from a working copy before you publish, to protect a source 553 cp capture/evidence.jpg publish/evidence.jpg 554 exiftool -all= -overwrite_original publish/evidence.jpg 555 shasum -a 256 publish/evidence.jpg # a different file, so a different hash: say so 556 ``` 557 558 Run this on the copy, after hashing the original, and never with `-overwrite_original` on anything 559 in the capture tree. Note the last block: a published, stripped file is a *different artefact* with a 560 different hash, and conflating the two is a straightforward way to be accused of altering evidence. 561 Record both hashes and the relationship between them. 562 563 What it does not tell you: metadata is written as easily as it is read, so a helpful 564 `CreateDate` is consistent-with and never proof. Most files pulled from platforms have had their 565 metadata stripped on upload, and that absence is normal rather than suspicious. 566 567 ## Case-management and pipeline tools 568 569 | Tool | What it does | 570 | --- | --- | 571 | [Auto Archiver](https://github.com/bellingcat/auto-archiver) | Bellingcat's pipeline: reads URLs from a spreadsheet, archives pages and media, hashes everything, writes results back. The right tool for continuous collection. | 572 | [Hunchly](https://www.hunch.ly/) | Captures every page you visit during a case automatically, with hashes and full-text search. Removes the discipline problem. | 573 | [Atlos](https://www.atlos.org/) | Collaborative platform for visual investigations with source tracking and review workflow. | 574 | [Lumen](https://lumendatabase.org/) | Archive of takedown notices — sometimes the only record that something existed. | 575 | [Web Archives](https://github.com/dessant/web-archives) | Browser extension that queries many archive services at once. | 576 577 ## Tool reference 578 579 | Tool | What it does | Cost | 580 | --- | --- | --- | 581 | [Archive.today](https://archive.today) | Archive any webpage and search for archived pages. | free | 582 | [Bellingcat TikTok Hashtag Analysis](https://github.com/bellingcat/tiktok-hashtag-analysis) | Archive content and metadata from TikTok posts that contain one or more specified hashtags | free | 583 | [Distill](https://distill.io/) | Distill is a website change monitoring tool that allows users to track changes on web pages. | partly free | 584 | [Wayback Machine](https://web.archive.org/) | The Wayback Machine is the Internet Archive's free tool for viewing and saving archived web pages, with over a trillion pages captured, widely used for… | free | 585 586 ## Pitfalls 587 588 - **Screenshots alone are weak.** No metadata, no hash, trivially edited. Capture the file. 589 - **Archives honour robots.txt and takedowns.** A site can retroactively remove itself from 590 Wayback. Your local copy is what survives. 591 - **Dynamic content does not archive well.** Infinite scroll, lazy loading and interactive maps 592 frequently fail; verify the capture actually shows what you saw. 593 - **Hash after acquisition, not after editing.** A hash of a file you already cropped proves 594 nothing about the original. 595 - **Archiving can notify.** Submitting a URL to a public archive creates a public record that 596 someone is interested in it. 597 - **A capture that succeeded is not a capture that worked.** A WARC full of 403s, a WACZ containing 598 a block page, and a `yt-dlp` channel pull truncated by rate limiting all exit zero. Read the CDX 599 status column, the crawl log and the row count before you file anything. 600 - **Old scripts hitting `GET /save/<url>` now get HTTP 429.** Anonymous Save Page Now is no longer 601 usable; a 429 body is still a body, so a script that does not check the status silently records a 602 failure as a success. Use the authenticated SPN2 POST. 603 - **A hash in your own notes is nearly worthless.** You could have written it at any time. The 604 third-party timestamp is what makes it an assertion about the past rather than about your memory. 605 606 ## Worked example 607 608 One datum: a URL to a post carrying a video clip, `https://example.com/@account/post/123`, cited in 609 a filing you have been asked to check. The URL, the hashes and the timestamps below are invented; 610 the order of the steps and what each one proves are not. 611 612 ```bash 613 # 1. what already exists, before you touch the page 614 curl -s 'https://web.archive.org/cdx/search/cdx?url=example.com/@account/post/123&output=json&fl=timestamp,digest,statuscode&collapse=digest' | jq -r '.[] | @tsv' 615 # timestamp digest statuscode 616 # 20260228183044 H4XQ... 200 617 # 20260302094112 ZK7B... 200 618 ``` 619 620 Two distinct payload digests, two days apart. Collapsing on digest is what revealed that: the page 621 was edited between 28 February and 2 March, which is a finding in itself and the reason to pull 622 both old captures before capturing the live page. 623 624 ```bash 625 # 2. local capture of the live page, byte-level, with the case in the warcinfo record 626 mkdir -p case-2026-014 && cd case-2026-014 627 wget --warc-file=live --warc-cdx --no-warc-compression \ 628 --warc-header="operator: J. Investigator" \ 629 --warc-header="description: case-2026-014, post cited at para 17" \ 630 --page-requisites --adjust-extension --no-parent \ 631 'https://example.com/@account/post/123' 632 awk 'NR>1 {print $5, $1}' live.cdx | sort | uniq -c 633 # 14 200 https://example.com/... 634 # 3 403 https://cdn.example.com/media/... 635 ``` 636 637 Three requisites returned 403, and they are the media files. The `wget` capture has the page text 638 and not the clip, which is the common outcome and the reason the next two steps exist rather than 639 being optional extras. 640 641 ```bash 642 # 3. the rendered page, in a browser, because the post body loads via JavaScript 643 docker run -v "$PWD/crawls:/crawls/" -it webrecorder/browsertrix-crawler crawl \ 644 --url 'https://example.com/@account/post/123' --scopeType page \ 645 --generateWACZ --text to-pages --collection case-2026-014 646 unzip -p crawls/collections/case-2026-014/case-2026-014.wacz datapackage-digest.json | jq . 647 # { "path": "datapackage.json", "hash": "sha256:9f2c...", "signedData": null } 648 ``` 649 650 `signedData: null` means the WACZ is unsigned — fine, because you are about to timestamp it 651 yourself. The digest over `datapackage.json` is the crawler's own integrity check on the archive; 652 record it, then stop relying on it and use your own hash. 653 654 ```bash 655 # 4. the media, with its provenance sidecar, which the WARC never got 656 yt-dlp --write-info-json --write-description --write-thumbnail --no-mtime \ 657 -o '%(upload_date)s_%(uploader_id)s_%(id)s.%(ext)s' \ 658 'https://example.com/@account/post/123' 659 jq -r '[.id,.uploader_id,.upload_date,.timestamp,.duration,.view_count] | @tsv' *.info.json 660 # 123 account 20260226 1772120000 47 18422 661 ``` 662 663 The sidecar says the clip was uploaded on 26 February — two days *before* the earliest Wayback 664 capture of the post, and before the edit the CDX diff exposed. That is the pivot: the video predates 665 the post text it now sits under, so the claim to check is whether the caption was changed, not 666 whether the footage is real. 667 668 ```bash 669 # 5. hash everything, in a stable order, before any further handling 670 cd .. && find ./case-2026-014 -type f -print0 | sort -z | xargs -0 shasum -a 256 > MANIFEST.sha256 671 wc -l MANIFEST.sha256 && shasum -a 256 MANIFEST.sha256 672 # 31 MANIFEST.sha256 673 # 4b1c8e... MANIFEST.sha256 674 ``` 675 676 ```bash 677 # 6. the step that makes the hash mean something: two independent timestamps 678 ots stamp MANIFEST.sha256 679 openssl ts -query -data MANIFEST.sha256 -sha256 -cert -out manifest.tsq 680 curl -s -H 'Content-Type: application/timestamp-query' \ 681 --data-binary @manifest.tsq https://freetsa.org/tsr -o manifest.tsr 682 openssl ts -reply -in manifest.tsr -text | grep -E 'Status:|Time stamp:' 683 # Status: Granted. 684 # Time stamp: Oct 4 04:06:07 2026 GMT 685 ``` 686 687 Now the claim is narrow and defensible: this set of 31 files, with these hashes, existed no later 688 than 04:06 UTC on 4 October 2026, attested by a TSA and by the Bitcoin blockchain, neither of which 689 you control. 690 691 ```bash 692 # 7. a third-party capture, submitted last so the local copy exists first 693 curl -s -X POST -H 'Accept: application/json' \ 694 -H "Authorization: LOW $IA_ACCESS_KEY:$IA_SECRET_KEY" \ 695 -d 'url=https://example.com/@account/post/123' -d 'capture_all=1' \ 696 'https://web.archive.org/save' | jq -r '.job_id' 697 # spn2-9f3c... 698 ``` 699 700 Then the same URL by hand into [archive.today](https://archive.today), and its short URL and UTC 701 capture time into `PROVENANCE.txt`. Submitting last is deliberate: both submissions are public, so 702 if the account notices and deletes the post, you already hold the capture. 703 704 What you can assert: the page as it stood at a timestamped moment, the clip with the 705 platform's own metadata, and a change history from Wayback showing the post was edited after the 706 video was uploaded. What it does not establish: that the footage shows what the caption claims, or 707 where or when it was filmed. Those are 708 [Image & Video Forensics](/sheets/osint/image-video-forensics) and 709 [Geolocation](/sheets/osint/geolocation) questions, and no amount of archiving substitutes for them. 710 711 What would falsify it: any manifest entry that fails verification, a timestamp token whose digest 712 does not match the manifest, or a WARC/WACZ log showing that the claimed resource failed and only 713 an error page was captured. The public archive copy is corroboration; the local capture, hashes 714 and independently verifiable timestamps carry the claim. 715 716 ## Broader catalogues 717 718 - [Archiving OSINT](https://tools.osintnewsletter.com/tool-categories/archiving-osint) 719 - [Data Extraction OSINT](https://tools.osintnewsletter.com/tool-categories/data-extraction-osint) 720 721 722 ## More tools 723 724 Further tools for this area from the OSINT Newsletter Tools Library ([Archiving OSINT](https://tools.osintnewsletter.com/tool-categories/archiving-osint)), excluding those already listed above. 725 726 | Tool | What it does | 727 | --- | --- | 728 | [4plebs](https://4plebs.org/) | A third-party archive and search service for historical 4chan content. | 729 | [4shared](https://www.4shared.com/) | File-sharing and cloud-storage platform with a public search function. | 730 | [Anna's Archive](https://annas-archive.gl/) | Shadow library search engine aggregating books, papers, and digital texts. | 731 | [Archivarix Tube Search](https://tube.archivarix.net/) | A free OSINT tool for discovering historical versions of YouTube videos using archived snapshots from archive data. | 732 | [Arctic Shift](https://arctic-shift.photon-reddit.com/) | A Reddit search and archival platform that enables investigators to search historical Reddit posts, comments, deleted content… | 733 | [deaditArchive](https://deaditarchive.netlify.app/) | A searchable archive of selected deleted or purged Reddit communities, preserving posts from communities that are no longer… | 734 | [Follow That Page](https://www.followthatpage.com/) | A web-monitoring service that tracks specified webpages and sends notifications when changes are detected. | 735 | [Free Full PDF](https://www.freefullpdf.com/) | Academic search engine focused on locating freely available full-text scientific PDFs, including journal articles, theses… | 736 | [GetProofAnchor](https://getproofanchor.com/) | A web-based tool that captures and preserves online content as verifiable digital evidence. | 737 | [Lookyloo](https://lookyloo.circl.lu/capture) | A web forensics tool that captures a webpage and maps every domain, resource, redirect, and third-party connection involved in… | 738 | [ProofSnap](https://getproofsnap.com/) | A browser extension that captures a live web page from the rendering browser and seals it as a verifiable archive. | 739 | [SearchShared](https://www.searchshared.info/) | File-sharing search engine for locating publicly indexed/shared files across multiple hosting services. | 740 741 ## Sources 742 743 Both catalogues below are maintained by other people and are considerably larger than 744 this page. Use them as the canonical index; this sheet is a working route through them. 745 746 - [Bellingcat's Online Investigation Toolkit](https://bellingcat.gitbook.io/toolkit) — ~340 tools, each with its own 747 review page covering cost, difficulty, requirements and limitations. 748 - [OSINT Newsletter Tools Library](https://tools.osintnewsletter.com) — ~280 tools, organised by investigative goal. 749 750 Neither publishes a licence, so nothing here is copied from them: tool names, one-line 751 descriptions, cost flags and links are catalogue facts, and the method and commentary are 752 this site's own. See [credits](/credits).