daemon-sec-cheatsheet

The cheatsheet vault for operators: AD, enumeration, exploitation, priv-esc, web, DFIR
git clone https://git.daemon-sec.xyz/daemon-sec-cheatsheet.git
Log | Files | Refs | README | LICENSE

archiving-and-evidence.md (40425B)


      1 ---
      2 title: "Archiving & Evidence Preservation"
      3 description: "Capture sources so they survive deletion, with the hashes and timestamps that make a capture defensible."
      4 category: osint
      5 subcategory: "Archiving & Analysis"
      6 tags: [osint, archiving, evidence, preservation]
      7 tools: [wget, browsertrix-crawler, yt-dlp, auto-archiver, wayback, archive-today, opentimestamps, openssl, exiftool, shasum]
      8 difficulty: intermediate
      9 updated: 2026-10-04
     10 references:
     11   - name: "Bellingcat's Online Investigation Toolkit"
     12     url: "https://bellingcat.gitbook.io/toolkit"
     13     author: "Bellingcat"
     14     license: none
     15     relation: derived
     16     note: "Tool catalogue: names, descriptions, cost flags and links for this area."
     17   - name: "OSINT Newsletter Tools Library"
     18     url: "https://tools.osintnewsletter.com"
     19     author: "The OSINT Newsletter"
     20     license: none
     21     relation: derived
     22     note: "Second tool catalogue, cross-checked against the above."
     23 ---
     24 
     25 ## What this covers
     26 
     27 Capturing a source so that it still exists after someone deletes it, and so that you can show the
     28 copy you hold is the copy you took. The capture itself is the easy half; the hash, the third-party
     29 timestamp and the written provenance are what make it worth anything to an editor, a lawyer or a
     30 court. Everything below is organised around one question: what does this step actually prove?
     31 
     32 ## Method
     33 
     34 1. **Check what already exists before you capture.** A Wayback snapshot from before the content
     35    became interesting is worth more than one you took after. Query the availability and CDX APIs
     36    first; if there is an older capture, that is your baseline and the diff against today is itself a
     37    finding.
     38 2. **Capture to something you do not control, and to something you do.** A third-party archive
     39    supplies a timestamp nobody can accuse you of forging. A local copy survives the archive being
     40    rate-limited, robots-excluded or taken down. Neither substitutes for the other.
     41 3. **Pick the capture tool by what the page is.** Static HTML and media: `wget --warc-file`.
     42    JavaScript-rendered, infinite-scroll or login-walled: a real browser driving the capture, which
     43    in practice means `browsertrix-crawler` producing WACZ. Platform video: `yt-dlp`, which gets you
     44    the metadata sidecar a screen recording never will.
     45 4. **Hash on acquisition, before you look at the file.** A hash taken after you cropped, converted
     46    or re-saved proves something about your working copy and nothing about the original.
     47 5. **Timestamp the manifest, not every file.** One hash over the list of hashes, submitted to
     48    OpenTimestamps or an RFC 3161 authority, binds the whole collection to a point in time at one
     49    cost. That is the step that converts "I have a hash" into "I had this hash on that date".
     50 6. **Write the provenance down in the same commit as the files.** URL, UTC time of capture, the
     51    tool and version, who handed it to you, and how you reached the page. The route to a finding is
     52    part of the finding, and it is the part you will forget first.
     53 7. **Get a second independent capture.** Two archives of the same URL taken by different
     54    infrastructure, or the same video from a second uploader, is the difference between a claim and
     55    a corroborated claim. Two copies that both trace to one upload are one copy.
     56 
     57 Judgement calls worth naming: submitting a URL to a public archive creates a public record that
     58 somebody is interested in it, so on a target that watches its referrers and its Wayback entries,
     59 capture locally first and submit later. And a capture of a page you reached while logged in
     60 contains your session — strip or redact before you share the WARC.
     61 
     62 ## Key tools
     63 
     64 ### wget (WARC capture)
     65 
     66 The standard way to get a byte-level record of an HTTP exchange rather than a rendered picture of
     67 it. A WARC stores the request and response headers alongside the body for every resource fetched,
     68 with a SHA-1 digest per record, so the file is a transcript of the conversation and not just its
     69 outcome. Use it whenever the page is server-rendered; it is the most portable evidence format in
     70 this area and every replay tool reads it.
     71 
     72 ```bash
     73 brew install wget                # or: apt install wget
     74 ```
     75 
     76 ```bash
     77 # the baseline capture: WARC transcript plus a CDX index of what went into it
     78 wget --warc-file=capture-001 --warc-cdx \
     79      --page-requisites --adjust-extension --no-parent \
     80      'https://example.com/page'
     81 
     82 # stamp the operator and case into the warcinfo record, so the file self-documents
     83 wget --warc-file=capture-002 --warc-cdx \
     84      --warc-header="operator: J. Investigator" \
     85      --warc-header="description: case-2026-014, post cited in filing" \
     86      --warc-header="robots: off" \
     87      --page-requisites --adjust-extension 'https://example.com/page'
     88 
     89 # leave it uncompressed while you are inspecting it; grep works on a plain WARC
     90 wget --warc-file=capture-003 --no-warc-compression 'https://example.com/page'
     91 
     92 # pull in the CDN that actually serves the images, or the capture renders blank
     93 wget --warc-file=capture-004 --warc-cdx \
     94      --page-requisites --convert-links --adjust-extension \
     95      --span-hosts --domains example.com,cdn.example.com \
     96      'https://example.com/page'
     97 
     98 # a whole section, politely: the delay is what stops you being blocked mid-capture
     99 wget --warc-file=capture-005 --warc-cdx --recursive --level=2 --no-parent \
    100      --wait=2 --random-wait --limit-rate=500k \
    101      -U 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0 Safari/537.36' \
    102      'https://example.com/newsroom/'
    103 
    104 # a list of URLs in one WARC, split at 1GB so the files stay movable
    105 wget --warc-file=capture-006 --warc-max-size=1G --warc-cdx -i urls.txt
    106 
    107 # second pass over the same site without re-storing what you already have
    108 wget --warc-file=capture-007 --warc-dedup=capture-005.cdx --recursive --no-parent \
    109      'https://example.com/newsroom/'
    110 
    111 # read back what the capture contains, without a replay tool
    112 zcat capture-001.warc.gz | grep -a '^WARC-Target-URI:' | sort -u
    113 awk 'NR>1 {print $5, $6, $1}' capture-001.cdx   # status, SHA-1 digest, URL
    114 ```
    115 
    116 Read the CDX file: the digest column is a base32 SHA-1 of the response payload, which gives you a
    117 free per-resource integrity check, and the status column tells you which requisites 404'd. A
    118 capture whose CDX is full of `302` and `403` rows did not get the page.
    119 
    120 What this does not prove: the WARC records what *your* client received from *an* IP at that moment.
    121 It does not prove the page looked that way to anyone else, does not survive geo-targeted or
    122 personalised content, and carries no third-party attestation of the time — the `WARC-Date` is your
    123 own clock. `--convert-links` rewrites the mirrored files on disk, not the WARC records, so the
    124 transcript stays pristine while the browsable copy is usable; keep both. Above all, `wget` executes
    125 no JavaScript, so on a modern single-page application it faithfully archives an empty shell. If the
    126 page needs a browser, use one.
    127 
    128 ### browsertrix-crawler
    129 
    130 A headless Chrome driven by Webrecorder's crawler, writing WARC and WACZ. This is the live,
    131 maintained answer for anything JavaScript-rendered, where `wget` returns a loading spinner. It
    132 scrolls, autoplays, runs site-specific behaviours for the big platforms, and extracts page text for
    133 search.
    134 
    135 The older Python crawlers in this niche have aged out: `wpull` carries a 2013–2016 copyright and no
    136 release since 2019, and `grab-site` still pins Python 3.7/3.8 and was dropped from nixpkgs after
    137 23.05. Neither is a reasonable thing to stand behind in 2026. Treat them as read-only history and
    138 use the crawler below.
    139 
    140 ```bash
    141 docker pull webrecorder/browsertrix-crawler
    142 ```
    143 
    144 ```bash
    145 # single page, as WACZ, with text extracted for full-text search
    146 docker run -v "$PWD/crawls:/crawls/" -it webrecorder/browsertrix-crawler crawl \
    147   --url 'https://example.com/post/123' --scopeType page \
    148   --generateWACZ --text to-pages --collection case-2026-014
    149 
    150 # the page plus whatever it links to on the same host, two levels deep
    151 docker run -v "$PWD/crawls:/crawls/" -it webrecorder/browsertrix-crawler crawl \
    152   --url 'https://example.com/newsroom/' --scopeType host --depth 2 \
    153   --generateWACZ --collection newsroom
    154 
    155 # a feed that only loads on scroll: give the behaviours time to finish
    156 docker run -v "$PWD/crawls:/crawls/" -it webrecorder/browsertrix-crawler crawl \
    157   --url 'https://example.com/feed' --scopeType page-spa \
    158   --behaviors autoscroll,autoplay,autofetch,siteSpecific \
    159   --behaviorTimeout 300 --generateWACZ --collection feed
    160 
    161 # several URLs from a file, one collection, so the manifest covers the set
    162 docker run -v "$PWD/crawls:/crawls/" -it webrecorder/browsertrix-crawler crawl \
    163   --urlFile /crawls/urls.txt --scopeType page --generateWACZ --collection batch-01
    164 
    165 # the WACZ is a zip: look inside before you trust it
    166 unzip -l crawls/collections/case-2026-014/case-2026-014.wacz
    167 unzip -p crawls/collections/case-2026-014/case-2026-014.wacz datapackage-digest.json | jq .
    168 
    169 # the crawl log is JSON Lines — this is where blocked and failed pages surface
    170 jq -r 'select(.logLevel=="error") | [.timestamp,.message] | @tsv' \
    171   crawls/collections/case-2026-014/logs/*.log
    172 ```
    173 
    174 A WACZ is a WARC plus an index, a page list and a signed digest manifest, and it opens in
    175 [ReplayWeb.page](https://replayweb.page/) with no server. That makes it the format to hand to
    176 someone who is not going to install anything. Check `datapackage-digest.json` after every crawl:
    177 if the crawler was blocked, you get a valid WACZ containing a block page, and nothing about the
    178 file itself says so.
    179 
    180 What it does not tell you: a browser capture is a recording of one rendering session, which means
    181 it includes whatever was personalised, A/B-tested or geo-fenced for that session. Running the same
    182 crawl from a different exit and getting different content is a finding about the site, not an error.
    183 Logged-in crawls bake the session into the archive — handle accordingly.
    184 
    185 ### yt-dlp
    186 
    187 Platform video and audio with its metadata intact. The point is not the video file; a screen
    188 recording gets you that. The point is the `.info.json` sidecar, which carries uploader ID, upload
    189 date, duration, view and comment counts, available formats and often the original title and
    190 description before anybody edited them.
    191 
    192 ```bash
    193 brew install yt-dlp              # or: pipx install yt-dlp
    194 ```
    195 
    196 ```bash
    197 # the provenance capture: video, metadata, description, thumbnail, subtitles
    198 yt-dlp --write-info-json --write-description --write-thumbnail \
    199        --write-subs --sub-langs 'all' --no-mtime \
    200        -o '%(upload_date)s_%(uploader_id)s_%(id)s.%(ext)s' URL
    201 
    202 # what formats exist, before you commit to one
    203 yt-dlp -F URL
    204 
    205 # the best single pre-merged file, so you archive bytes the platform served
    206 yt-dlp -f 'best[ext=mp4]/best' --no-mtime URL
    207 
    208 # metadata only — enough to date and attribute a clip without downloading it
    209 yt-dlp --skip-download --write-info-json --write-thumbnail URL
    210 
    211 # comments too; they routinely date an event more precisely than the upload does
    212 yt-dlp --write-info-json --write-comments --skip-download URL
    213 
    214 # a whole channel, newest first, with an archive file so re-runs are incremental
    215 yt-dlp --write-info-json --no-mtime --download-archive seen.txt \
    216        -o '%(upload_date)s_%(id)s.%(ext)s' 'https://example.com/@account/videos'
    217 
    218 # the intermediate pages the extractor fetched — useful when an extractor misreads a page
    219 yt-dlp --write-pages --skip-download URL
    220 
    221 # the fields you will actually quote, straight out of the sidecar
    222 jq -r '[.id,.uploader_id,.upload_date,.timestamp,.duration,.view_count,.title] | @tsv' *.info.json
    223 ```
    224 
    225 `--no-mtime` matters: without it the file's mtime is set from the upload date, which silently
    226 overwrites your own acquisition timeline with the platform's claim. `timestamp` in the sidecar is
    227 Unix epoch seconds and is usually more precise than `upload_date`, which is date-only and rendered
    228 in the platform's timezone of choice.
    229 
    230 What the sidecar does not establish: every field in it is the platform's assertion, repeated.
    231 `upload_date` is when it was uploaded, never when it was filmed, and a re-upload resets it.
    232 Extractors break when platforms change, so a field that is `null` may mean absent or may mean the
    233 extractor lost it — check against the page. Rate limits and age or region gates will silently
    234 truncate a channel pull; compare the count you got against the count the channel claims.
    235 
    236 ### Wayback Machine, from the command line
    237 
    238 Three separate endpoints, used at three different moments. Availability answers "is there already a
    239 capture"; CDX answers "what captures exist, and did the content change between them"; Save Page Now
    240 creates a new one.
    241 
    242 ```bash
    243 # is there a capture at all, and what is the nearest one to a date
    244 curl -s 'https://archive.org/wayback/available?url=example.com/page&timestamp=20240101' | jq .
    245 
    246 # every capture of a URL, with the payload digest that reveals real edits
    247 curl -s 'https://web.archive.org/cdx/search/cdx?url=example.com/page&output=json&fl=timestamp,original,statuscode,digest,length' | jq -r '.[] | @tsv'
    248 
    249 # collapse consecutive identical captures: what is left is the change history
    250 curl -s 'https://web.archive.org/cdx/search/cdx?url=example.com/page&output=json&fl=timestamp,digest&collapse=digest' | jq -r '.[] | @tsv'
    251 
    252 # everything ever captured under a host, which is how you find deleted pages
    253 curl -s 'https://web.archive.org/cdx/search/cdx?url=example.com&matchType=domain&output=json&fl=timestamp,original,statuscode&limit=500' | jq -r '.[] | @tsv'
    254 
    255 # only the captures in a window, for a page that mattered on one specific day
    256 curl -s 'https://web.archive.org/cdx/search/cdx?url=example.com/page&from=20260301&to=20260331&output=json&fl=timestamp,digest' | jq -r '.[] | @tsv'
    257 
    258 # submit a new capture: SPN2, authenticated with Internet Archive S3-style keys
    259 curl -s -X POST -H 'Accept: application/json' \
    260      -H "Authorization: LOW $IA_ACCESS_KEY:$IA_SECRET_KEY" \
    261      -d 'url=https://example.com/page' -d 'capture_all=1' -d 'capture_screenshot=1' \
    262      'https://web.archive.org/save' | jq .
    263 
    264 # poll the job until it reports success, and keep the returned timestamp
    265 curl -s -H "Authorization: LOW $IA_ACCESS_KEY:$IA_SECRET_KEY" \
    266      "https://web.archive.org/save/status/$JOB_ID" | jq '{status, timestamp, original_url}'
    267 
    268 # fetch the archived copy itself, and hash what you got
    269 curl -sL 'https://web.archive.org/web/20260301120000id_/https://example.com/page' | shasum -a 256
    270 ```
    271 
    272 Get the key pair from [archive.org/account/s3.php](https://archive.org/account/s3.php) and export
    273 it as `IA_ACCESS_KEY` / `IA_SECRET_KEY`. **The unauthenticated `GET /save/<url>` form is no longer
    274 usable for work** — an anonymous request to it returns HTTP 429 rather than a capture. Scripts that
    275 still rely on it fail silently, because a 429 body looks like a page.
    276 
    277 The `id_` infix in that last URL asks for the original bytes without Wayback's navigation banner
    278 injected, which is what you want if you intend to hash or diff the archived copy. The `digest`
    279 column in CDX output is the payload hash: two captures sharing a digest are byte-identical, so
    280 collapsing on it turns a thousand snapshots into the handful of moments the page actually changed.
    281 That is the single most useful thing the CDX API does.
    282 
    283 What Wayback does not give you: permanence. A site can retroactively exclude itself, and captures
    284 do disappear. It also honours robots directives at crawl time, so absence of a capture is not
    285 evidence the page did not exist. And a Save Page Now capture is still just a capture — it attests
    286 that the Internet Archive saw that content at that time, which is a strong third-party claim about
    287 time and a weak one about authenticity.
    288 
    289 ### archive.today
    290 
    291 A separate archive with a separate failure mode, which is exactly why you use it alongside Wayback.
    292 It renders pages that Wayback flattens, ignores robots directives, and has a long record of not
    293 removing things. There is no API and the submission form sits behind bot protection, so this one is
    294 a browser job — do not script it.
    295 
    296 ```text
    297 1. Open https://archive.today  (archive.ph and archive.is are the same service;
    298    if one mirror is blocked where you are, try another)
    299 2. Paste the URL into the lower box, "My url is alive and I want to archive its content"
    300 3. Solve the challenge if you get one. Do not automate this step — it is the step
    301    that gets the service to block your address.
    302 4. Wait for the capture. The result URL looks like https://archive.ph/AbC12
    303    and the page header shows the capture time in UTC.
    304 5. Click "screenshot" in the header to get the full-page render as a separate
    305    artefact, and save it. The HTML capture and the screenshot fail differently.
    306 6. Record the short URL AND the capture time. The short URL is your citation.
    307 7. Before submitting, use the upper box to search for existing captures of the
    308    same URL — the same reason you query Wayback's CDX API first.
    309 ```
    310 
    311 Two things to know. The service is deliberately opaque about its infrastructure and funding, which
    312 means you should not treat it as your only copy of anything; keep the local WARC. And it fetches
    313 the page itself, from its own address, so a submission does not leak your IP to the target — but it
    314 does create a public, searchable record that someone archived that URL at that minute.
    315 
    316 ### Hashing and the manifest
    317 
    318 Hashes are the whole evidentiary argument. A hash recorded at acquisition and published or
    319 timestamped separately lets you show, later, that the file you are producing is the file you took.
    320 The manifest pattern — one file listing every hash, then one hash of the manifest — scales that to a
    321 collection without timestamping a thousand files.
    322 
    323 ```bash
    324 # macOS ships shasum; coreutils ships sha256sum. Same output format.
    325 shasum -a 256 evidence.mp4
    326 sha256sum evidence.mp4
    327 ```
    328 
    329 ```bash
    330 # hash the whole capture tree in a stable order, so the manifest is reproducible
    331 find ./capture -type f -print0 | sort -z | xargs -0 shasum -a 256 > MANIFEST.sha256
    332 
    333 # the manifest's own hash: this single line is what you timestamp and publish
    334 shasum -a 256 MANIFEST.sha256 | tee MANIFEST.sha256.sha256
    335 
    336 # verify the collection, any time later
    337 shasum -a 256 -c MANIFEST.sha256
    338 
    339 # show only what failed, which is what you actually want on a large set
    340 shasum -a 256 -c MANIFEST.sha256 2>/dev/null | grep -v ': OK$'
    341 
    342 # a provenance record that travels with the files, in the same directory
    343 cat > PROVENANCE.txt <<'EOF'
    344 case:        2026-014
    345 url:         https://example.com/post/123
    346 captured_at: 2026-10-04T04:07:56Z   # UTC, from `date -u +%FT%TZ`
    347 captured_by: J. Investigator
    348 tooling:     wget 1.25.0 --warc-file; yt-dlp 2026.08.19
    349 route:       linked from https://example.com/newsroom/ on 2026-10-03
    350 notes:       page required no login; no personalisation observed
    351 EOF
    352 
    353 # fold the provenance into the manifest so the two cannot drift apart
    354 shasum -a 256 PROVENANCE.txt >> MANIFEST.sha256
    355 
    356 # find duplicates across two collections by hash, not by filename
    357 sort MANIFEST.sha256 other/MANIFEST.sha256 | awk '{print $1}' | sort | uniq -d
    358 ```
    359 
    360 Use SHA-256. MD5 and SHA-1 still appear in forensic tooling and are fine as identifiers, but a
    361 collision-prone hash invites an argument you do not need to have. Note that `find | sort` matters:
    362 an unsorted manifest changes order between runs and filesystems, which makes two manifests of the
    363 same files look different.
    364 
    365 What a hash proves and does not: it proves the bytes have not changed since the hash was taken. It
    366 says nothing about when they were taken, nothing about where they came from, and nothing about
    367 whether the content is true. On its own a hash in your own notes is worth very little, because you
    368 could have written it at any time — which is what the next section fixes.
    369 
    370 ### Timestamping: OpenTimestamps and RFC 3161
    371 
    372 The step that turns a hash into evidence. A timestamp is a third party's signed assertion that a
    373 given digest existed before a given moment. Without it, your hash only proves internal
    374 consistency; with it, you can show you held that exact file on that date and could not have
    375 produced it later.
    376 
    377 ```bash
    378 pipx install opentimestamps-client     # the `ots` command
    379 # openssl is already present for the RFC 3161 route
    380 ```
    381 
    382 ```bash
    383 # OpenTimestamps: anchors the digest in the Bitcoin blockchain, free, no account
    384 ots stamp MANIFEST.sha256
    385 
    386 # what the proof currently contains: pending calendar attestations, then a block
    387 ots info MANIFEST.sha256.ots
    388 
    389 # calendars need an hour or so to get into a block; upgrade the proof afterwards
    390 ots upgrade MANIFEST.sha256.ots
    391 
    392 # verify. Full verification wants a local Bitcoin node; without one you are
    393 # trusting the calendar, which is weaker but still a third party
    394 ots verify MANIFEST.sha256.ots
    395 
    396 # verify a digest you were given, without holding the file
    397 ots verify -d 355c357708e4840d12b0d4284d9fe1a911c79c013953874f03068b81127f3363 MANIFEST.sha256.ots
    398 
    399 # RFC 3161 route: build a timestamp query over the file's SHA-256
    400 openssl ts -query -data MANIFEST.sha256 -sha256 -cert -out manifest.tsq
    401 
    402 # send it to a Time Stamping Authority and keep the signed reply
    403 curl -s -H 'Content-Type: application/timestamp-query' \
    404      --data-binary @manifest.tsq https://freetsa.org/tsr -o manifest.tsr
    405 
    406 # read the token: the TSA's time, its identity and the digest it signed over
    407 openssl ts -reply -in manifest.tsr -text
    408 
    409 # verify the token against the TSA's chain — this is the check a reviewer repeats
    410 curl -sO https://freetsa.org/files/cacert.pem
    411 curl -sO https://freetsa.org/files/tsa.crt
    412 openssl ts -verify -data MANIFEST.sha256 -in manifest.tsr \
    413            -CAfile cacert.pem -untrusted tsa.crt
    414 ```
    415 
    416 A successful verification prints `Verification: OK`, and that line is the thing you cite. Keep the
    417 `.tsr`, the CA chain you verified against, and the exact file — the token is signed over the digest,
    418 so a reviewer needs the file to recompute it.
    419 
    420 Which route to use: OpenTimestamps costs nothing, needs no account, and anchors to a public
    421 blockchain, so the proof outlives any single company — but it is coarse (block granularity, roughly
    422 an hour) and full verification wants a Bitcoin node. An RFC 3161 TSA gives you a precise, signed,
    423 PKI-rooted token that institutional reviewers already recognise, at the price of trusting that TSA
    424 and its certificate remaining validatable. Do both; they are cheap and they fail differently.
    425 
    426 What a timestamp does not prove: that the content is authentic, that the capture was complete, or
    427 that you obtained it lawfully. It proves the file existed no later than that moment. That is a
    428 narrow claim, and stating it narrowly is what makes it credible.
    429 
    430 ### Bellingcat Auto Archiver
    431 
    432 The pipeline to set up when you are collecting continuously rather than case by case. It reads URLs
    433 from a feeder, runs a platform-specific extractor, then runs enrichers that hash, thumbnail,
    434 timestamp and push to Wayback, and writes the results to a database you can read as a case index.
    435 It is the only tool here that performs the hash-and-timestamp discipline on every item without you
    436 remembering to.
    437 
    438 ```bash
    439 pipx install auto-archiver
    440 auto-archiver --help           # the flag list is generated from the installed modules
    441 ```
    442 
    443 Configuration is a single YAML file; the default path is `secrets/orchestration.yaml`. Only modules
    444 named under `steps` are loaded, so the file is both configuration and a declaration of the pipeline.
    445 
    446 ```yaml
    447 # orchestration.yaml
    448 steps:
    449   feeders: [cli_feeder]
    450   extractors: [generic_extractor]
    451   enrichers: [hash_enricher, meta_enricher, thumbnail_enricher,
    452               timestamping_enricher, opentimestamps_enricher]
    453   databases: [csv_db, console_db]
    454   storages: [local_storage]
    455   formatters: [html_formatter]
    456 
    457 hash_enricher:
    458   algorithm: SHA-256          # the only alternative is SHA3-512
    459 
    460 timestamping_enricher:
    461   tsa_urls:
    462     - http://timestamp.identrust.com
    463     - http://zeitstempel.dfn.de
    464   allow_selfsigned: false     # leaving this false is the whole point of the module
    465 
    466 local_storage:
    467   save_to: ./archived
    468   path_generator: url         # flat | url | random
    469   filename_generator: static  # names the file after its hash
    470 
    471 csv_db:
    472   csv_file: ./archived/results.csv
    473 
    474 logging:
    475   level: INFO
    476   file: ./archived/archiver.log
    477 ```
    478 
    479 ```bash
    480 # one URL through the whole pipeline
    481 auto-archiver --config orchestration.yaml 'https://example.com/post/123'
    482 
    483 # several at once; the CSV gets one row per URL
    484 auto-archiver --config orchestration.yaml 'https://example.com/a' 'https://example.com/b'
    485 
    486 # no config file at all, for a one-off capture with sane defaults
    487 auto-archiver --mode simple 'https://example.com/post/123'
    488 
    489 # override the enricher list without editing the file
    490 auto-archiver --config orchestration.yaml \
    491   --enrichers hash_enricher timestamping_enricher opentimestamps_enricher \
    492   'https://example.com/post/123'
    493 
    494 # push to Wayback as part of the run, with the same Internet Archive key pair
    495 auto-archiver --config orchestration.yaml \
    496   --enrichers hash_enricher wayback_extractor_enricher \
    497   --wayback_extractor_enricher.key "$IA_ACCESS_KEY" \
    498   --wayback_extractor_enricher.secret "$IA_SECRET_KEY" \
    499   'https://example.com/post/123'
    500 
    501 # feed a column of URLs from a spreadsheet export
    502 auto-archiver --config orchestration.yaml \
    503   --feeders csv_feeder --csv_feeder.files urls.csv --csv_feeder.column link
    504 
    505 # turn up logging when an extractor silently returns nothing
    506 auto-archiver --config orchestration.yaml --logging.level DEBUG 'https://example.com/post/123'
    507 
    508 # the results CSV is the case index: one row per URL, with hashes and archive links
    509 column -s, -t < ./archived/results.csv | less -S
    510 ```
    511 
    512 Two enrichers do the evidentiary work and they are not interchangeable. `timestamping_enricher`
    513 aggregates the item's file hashes into one text file and gets an RFC 3161 token over it from each
    514 TSA in the list; `opentimestamps_enricher` anchors the same digest via the OpenTimestamps
    515 calendars. Note that `freetsa.org` is deliberately absent from the shipped TSA defaults because its
    516 certificate does not validate against system authorities — you can add it, but only together with
    517 `allow_selfsigned: true`, which weakens exactly the property you wanted.
    518 
    519 The `gsheet_feeder_db` module is what makes this a team tool: analysts paste URLs into a Google
    520 Sheet, the archiver writes the hash, the stored path and the archive links back into adjacent
    521 columns. The `wacz_extractor_enricher` shells out to Docker for browser-based capture, and
    522 `wayback_extractor_enricher` needs the same key pair as Save Page Now.
    523 
    524 Watch the version. The client checks on startup and says so when it is behind, and per-platform
    525 extractors break often enough that an old release produces empty captures that look like
    526 successes. After every run, scan the results CSV for rows that carry a hash but no media.
    527 
    528 ### ExifTool, for capture-time metadata
    529 
    530 Covered in depth on [Image & Video Forensics](/sheets/osint/image-video-forensics); the archiving
    531 use is narrower. You are not asking whether the file is manipulated, you are recording what the
    532 file asserted about itself at the moment it entered your custody, so that a later disagreement is
    533 about the metadata rather than about your handling.
    534 
    535 ```bash
    536 # the full tag dump, archived next to the hash and never edited again
    537 exiftool -j -g1 -a -u capture/evidence.mp4 > capture/evidence.mp4.exif.json
    538 
    539 # the fields that date and attribute a capture, in one line
    540 exiftool -CreateDate -ModifyDate -FileModifyDate -GPSPosition -Make -Model -Software \
    541          capture/evidence.mp4
    542 
    543 # every timestamp in the file, grouped by where it came from
    544 exiftool -time:all -a -G1 -s capture/evidence.mp4
    545 
    546 # the whole capture tree as one CSV, which doubles as a collection inventory
    547 exiftool -r -csv -FileName -FileSize -MIMEType -CreateDate -GPSPosition ./capture/ > inventory.csv
    548 
    549 # flag files whose structure does not match their declared format
    550 exiftool -r -validate -warning -a ./capture/
    551 
    552 # strip metadata from a working copy before you publish, to protect a source
    553 cp capture/evidence.jpg publish/evidence.jpg
    554 exiftool -all= -overwrite_original publish/evidence.jpg
    555 shasum -a 256 publish/evidence.jpg    # a different file, so a different hash: say so
    556 ```
    557 
    558 Run this on the copy, after hashing the original, and never with `-overwrite_original` on anything
    559 in the capture tree. Note the last block: a published, stripped file is a *different artefact* with a
    560 different hash, and conflating the two is a straightforward way to be accused of altering evidence.
    561 Record both hashes and the relationship between them.
    562 
    563 What it does not tell you: metadata is written as easily as it is read, so a helpful
    564 `CreateDate` is consistent-with and never proof. Most files pulled from platforms have had their
    565 metadata stripped on upload, and that absence is normal rather than suspicious.
    566 
    567 ## Case-management and pipeline tools
    568 
    569 | Tool | What it does |
    570 | --- | --- |
    571 | [Auto Archiver](https://github.com/bellingcat/auto-archiver) | Bellingcat's pipeline: reads URLs from a spreadsheet, archives pages and media, hashes everything, writes results back. The right tool for continuous collection. |
    572 | [Hunchly](https://www.hunch.ly/) | Captures every page you visit during a case automatically, with hashes and full-text search. Removes the discipline problem. |
    573 | [Atlos](https://www.atlos.org/) | Collaborative platform for visual investigations with source tracking and review workflow. |
    574 | [Lumen](https://lumendatabase.org/) | Archive of takedown notices — sometimes the only record that something existed. |
    575 | [Web Archives](https://github.com/dessant/web-archives) | Browser extension that queries many archive services at once. |
    576 
    577 ## Tool reference
    578 
    579 | Tool | What it does | Cost |
    580 | --- | --- | --- |
    581 | [Archive.today](https://archive.today) | Archive any webpage and search for archived pages. | free |
    582 | [Bellingcat TikTok Hashtag Analysis](https://github.com/bellingcat/tiktok-hashtag-analysis) | Archive content and metadata from TikTok posts that contain one or more specified hashtags | free |
    583 | [Distill](https://distill.io/) | Distill is a website change monitoring tool that allows users to track changes on web pages. | partly free |
    584 | [Wayback Machine](https://web.archive.org/) | The Wayback Machine is the Internet Archive's free tool for viewing and saving archived web pages, with over a trillion pages captured, widely used for… | free |
    585 
    586 ## Pitfalls
    587 
    588 - **Screenshots alone are weak.** No metadata, no hash, trivially edited. Capture the file.
    589 - **Archives honour robots.txt and takedowns.** A site can retroactively remove itself from
    590   Wayback. Your local copy is what survives.
    591 - **Dynamic content does not archive well.** Infinite scroll, lazy loading and interactive maps
    592   frequently fail; verify the capture actually shows what you saw.
    593 - **Hash after acquisition, not after editing.** A hash of a file you already cropped proves
    594   nothing about the original.
    595 - **Archiving can notify.** Submitting a URL to a public archive creates a public record that
    596   someone is interested in it.
    597 - **A capture that succeeded is not a capture that worked.** A WARC full of 403s, a WACZ containing
    598   a block page, and a `yt-dlp` channel pull truncated by rate limiting all exit zero. Read the CDX
    599   status column, the crawl log and the row count before you file anything.
    600 - **Old scripts hitting `GET /save/<url>` now get HTTP 429.** Anonymous Save Page Now is no longer
    601   usable; a 429 body is still a body, so a script that does not check the status silently records a
    602   failure as a success. Use the authenticated SPN2 POST.
    603 - **A hash in your own notes is nearly worthless.** You could have written it at any time. The
    604   third-party timestamp is what makes it an assertion about the past rather than about your memory.
    605 
    606 ## Worked example
    607 
    608 One datum: a URL to a post carrying a video clip, `https://example.com/@account/post/123`, cited in
    609 a filing you have been asked to check. The URL, the hashes and the timestamps below are invented;
    610 the order of the steps and what each one proves are not.
    611 
    612 ```bash
    613 # 1. what already exists, before you touch the page
    614 curl -s 'https://web.archive.org/cdx/search/cdx?url=example.com/@account/post/123&output=json&fl=timestamp,digest,statuscode&collapse=digest' | jq -r '.[] | @tsv'
    615 # timestamp        digest                            statuscode
    616 # 20260228183044   H4XQ...                           200
    617 # 20260302094112   ZK7B...                           200
    618 ```
    619 
    620 Two distinct payload digests, two days apart. Collapsing on digest is what revealed that: the page
    621 was edited between 28 February and 2 March, which is a finding in itself and the reason to pull
    622 both old captures before capturing the live page.
    623 
    624 ```bash
    625 # 2. local capture of the live page, byte-level, with the case in the warcinfo record
    626 mkdir -p case-2026-014 && cd case-2026-014
    627 wget --warc-file=live --warc-cdx --no-warc-compression \
    628      --warc-header="operator: J. Investigator" \
    629      --warc-header="description: case-2026-014, post cited at para 17" \
    630      --page-requisites --adjust-extension --no-parent \
    631      'https://example.com/@account/post/123'
    632 awk 'NR>1 {print $5, $1}' live.cdx | sort | uniq -c
    633 #  14 200 https://example.com/...
    634 #   3 403 https://cdn.example.com/media/...
    635 ```
    636 
    637 Three requisites returned 403, and they are the media files. The `wget` capture has the page text
    638 and not the clip, which is the common outcome and the reason the next two steps exist rather than
    639 being optional extras.
    640 
    641 ```bash
    642 # 3. the rendered page, in a browser, because the post body loads via JavaScript
    643 docker run -v "$PWD/crawls:/crawls/" -it webrecorder/browsertrix-crawler crawl \
    644   --url 'https://example.com/@account/post/123' --scopeType page \
    645   --generateWACZ --text to-pages --collection case-2026-014
    646 unzip -p crawls/collections/case-2026-014/case-2026-014.wacz datapackage-digest.json | jq .
    647 # { "path": "datapackage.json", "hash": "sha256:9f2c...", "signedData": null }
    648 ```
    649 
    650 `signedData: null` means the WACZ is unsigned — fine, because you are about to timestamp it
    651 yourself. The digest over `datapackage.json` is the crawler's own integrity check on the archive;
    652 record it, then stop relying on it and use your own hash.
    653 
    654 ```bash
    655 # 4. the media, with its provenance sidecar, which the WARC never got
    656 yt-dlp --write-info-json --write-description --write-thumbnail --no-mtime \
    657        -o '%(upload_date)s_%(uploader_id)s_%(id)s.%(ext)s' \
    658        'https://example.com/@account/post/123'
    659 jq -r '[.id,.uploader_id,.upload_date,.timestamp,.duration,.view_count] | @tsv' *.info.json
    660 # 123   account   20260226   1772120000   47   18422
    661 ```
    662 
    663 The sidecar says the clip was uploaded on 26 February — two days *before* the earliest Wayback
    664 capture of the post, and before the edit the CDX diff exposed. That is the pivot: the video predates
    665 the post text it now sits under, so the claim to check is whether the caption was changed, not
    666 whether the footage is real.
    667 
    668 ```bash
    669 # 5. hash everything, in a stable order, before any further handling
    670 cd .. && find ./case-2026-014 -type f -print0 | sort -z | xargs -0 shasum -a 256 > MANIFEST.sha256
    671 wc -l MANIFEST.sha256 && shasum -a 256 MANIFEST.sha256
    672 # 31 MANIFEST.sha256
    673 # 4b1c8e... MANIFEST.sha256
    674 ```
    675 
    676 ```bash
    677 # 6. the step that makes the hash mean something: two independent timestamps
    678 ots stamp MANIFEST.sha256
    679 openssl ts -query -data MANIFEST.sha256 -sha256 -cert -out manifest.tsq
    680 curl -s -H 'Content-Type: application/timestamp-query' \
    681      --data-binary @manifest.tsq https://freetsa.org/tsr -o manifest.tsr
    682 openssl ts -reply -in manifest.tsr -text | grep -E 'Status:|Time stamp:'
    683 # Status: Granted.
    684 # Time stamp: Oct  4 04:06:07 2026 GMT
    685 ```
    686 
    687 Now the claim is narrow and defensible: this set of 31 files, with these hashes, existed no later
    688 than 04:06 UTC on 4 October 2026, attested by a TSA and by the Bitcoin blockchain, neither of which
    689 you control.
    690 
    691 ```bash
    692 # 7. a third-party capture, submitted last so the local copy exists first
    693 curl -s -X POST -H 'Accept: application/json' \
    694      -H "Authorization: LOW $IA_ACCESS_KEY:$IA_SECRET_KEY" \
    695      -d 'url=https://example.com/@account/post/123' -d 'capture_all=1' \
    696      'https://web.archive.org/save' | jq -r '.job_id'
    697 # spn2-9f3c...
    698 ```
    699 
    700 Then the same URL by hand into [archive.today](https://archive.today), and its short URL and UTC
    701 capture time into `PROVENANCE.txt`. Submitting last is deliberate: both submissions are public, so
    702 if the account notices and deletes the post, you already hold the capture.
    703 
    704 What you can assert: the page as it stood at a timestamped moment, the clip with the
    705 platform's own metadata, and a change history from Wayback showing the post was edited after the
    706 video was uploaded. What it does not establish: that the footage shows what the caption claims, or
    707 where or when it was filmed. Those are
    708 [Image & Video Forensics](/sheets/osint/image-video-forensics) and
    709 [Geolocation](/sheets/osint/geolocation) questions, and no amount of archiving substitutes for them.
    710 
    711 What would falsify it: any manifest entry that fails verification, a timestamp token whose digest
    712 does not match the manifest, or a WARC/WACZ log showing that the claimed resource failed and only
    713 an error page was captured. The public archive copy is corroboration; the local capture, hashes
    714 and independently verifiable timestamps carry the claim.
    715 
    716 ## Broader catalogues
    717 
    718 - [Archiving OSINT](https://tools.osintnewsletter.com/tool-categories/archiving-osint)
    719 - [Data Extraction OSINT](https://tools.osintnewsletter.com/tool-categories/data-extraction-osint)
    720 
    721 
    722 ## More tools
    723 
    724 Further tools for this area from the OSINT Newsletter Tools Library ([Archiving OSINT](https://tools.osintnewsletter.com/tool-categories/archiving-osint)), excluding those already listed above.
    725 
    726 | Tool | What it does |
    727 | --- | --- |
    728 | [4plebs](https://4plebs.org/) | A third-party archive and search service for historical 4chan content. |
    729 | [4shared](https://www.4shared.com/) | File-sharing and cloud-storage platform with a public search function. |
    730 | [Anna's Archive](https://annas-archive.gl/) | Shadow library search engine aggregating books, papers, and digital texts. |
    731 | [Archivarix Tube Search](https://tube.archivarix.net/) | A free OSINT tool for discovering historical versions of YouTube videos using archived snapshots from archive data. |
    732 | [Arctic Shift](https://arctic-shift.photon-reddit.com/) | A Reddit search and archival platform that enables investigators to search historical Reddit posts, comments, deleted content… |
    733 | [deaditArchive](https://deaditarchive.netlify.app/) | A searchable archive of selected deleted or purged Reddit communities, preserving posts from communities that are no longer… |
    734 | [Follow That Page](https://www.followthatpage.com/) | A web-monitoring service that tracks specified webpages and sends notifications when changes are detected. |
    735 | [Free Full PDF](https://www.freefullpdf.com/) | Academic search engine focused on locating freely available full-text scientific PDFs, including journal articles, theses… |
    736 | [GetProofAnchor](https://getproofanchor.com/) | A web-based tool that captures and preserves online content as verifiable digital evidence. |
    737 | [Lookyloo](https://lookyloo.circl.lu/capture) | A web forensics tool that captures a webpage and maps every domain, resource, redirect, and third-party connection involved in… |
    738 | [ProofSnap](https://getproofsnap.com/) | A browser extension that captures a live web page from the rendering browser and seals it as a verifiable archive. |
    739 | [SearchShared](https://www.searchshared.info/) | File-sharing search engine for locating publicly indexed/shared files across multiple hosting services. |
    740 
    741 ## Sources
    742 
    743 Both catalogues below are maintained by other people and are considerably larger than
    744 this page. Use them as the canonical index; this sheet is a working route through them.
    745 
    746 - [Bellingcat's Online Investigation Toolkit](https://bellingcat.gitbook.io/toolkit) — ~340 tools, each with its own
    747   review page covering cost, difficulty, requirements and limitations.
    748 - [OSINT Newsletter Tools Library](https://tools.osintnewsletter.com) — ~280 tools, organised by investigative goal.
    749 
    750 Neither publishes a licence, so nothing here is copied from them: tool names, one-line
    751 descriptions, cost flags and links are catalogue facts, and the method and commentary are
    752 this site's own. See [credits](/credits).