internet-archival-guide.md (25781B)
1 --- 2 title: "Internet Archival Guide" 3 description: "sudo curl -L https://github.com/yt-dlp/yt-dlp/releases/latest/download/yt-dlp -o /usr/local/bin/yt-dlp sudo chmod a+rx /usr/local/bin/yt-dlp" 4 category: tools 5 tags: ["tools", "adcs"] 6 tools: [] 7 difficulty: intermediate 8 updated: "2026-08-10" 9 source: "vault:Tools/Internet-Archival-Guide/Internet Archival Guide.md" 10 --- 11 # Linux 12 sudo curl -L https://github.com/yt-dlp/yt-dlp/releases/latest/download/yt-dlp -o /usr/local/bin/yt-dlp 13 sudo chmod a+rx /usr/local/bin/yt-dlp 14 15 # Mac 16 brew install yt-dlp 17 18 # Keep updated (run regularly) 19 yt-dlp -U 20 ``` 21 22 > [!important]+ ffmpeg is required 23 > `fas:TriangleExclamation` 24 > yt-dlp uses ffmpeg to merge video and audio streams. Install from [ffmpeg.org](https://ffmpeg.org/download.html). On Windows, place `ffmpeg.exe` in the **same folder** as `yt-dlp.exe`. 25 26 **Core yt-dlp commands:** 27 28 ```bash 29 # Download best quality (auto format selection) 30 yt-dlp https://www.youtube.com/watch?v=VIDEOID 31 32 # List all available formats first 33 yt-dlp -F https://www.youtube.com/watch?v=VIDEOID 34 35 # Download best MP4 with merged audio+video (most compatible) 36 yt-dlp -f "bestvideo[ext=mp4]+bestaudio[ext=m4a]/best[ext=mp4]/best" https://www.youtube.com/watch?v=VIDEOID 37 38 # Download with metadata, thumbnail, and subtitles 39 yt-dlp --write-description --write-info-json --write-thumbnail --write-subs https://www.youtube.com/watch?v=VIDEOID 40 41 # Download audio only as MP3 42 yt-dlp -x --audio-format mp3 https://www.youtube.com/watch?v=VIDEOID 43 44 # Download entire channel (skip already-downloaded, continue on errors) 45 yt-dlp --ignore-errors --continue https://www.youtube.com/c/CHANNELNAME 46 47 # Download entire channel with full metadata + organised output 48 yt-dlp --ignore-errors --continue \ 49 --write-description --write-info-json \ 50 --write-thumbnail --write-subs \ 51 --output "%(uploader)s/%(upload_date)s-%(title)s.%(ext)s" \ 52 https://www.youtube.com/c/CHANNELNAME 53 54 # Download autogenerated subtitles (great for text-searching streams) 55 yt-dlp --write-auto-subs --sub-format vtt https://www.youtube.com/watch?v=VIDEOID 56 57 # Download a specific time segment only 58 yt-dlp --download-sections "*01:30-05:45" https://www.youtube.com/watch?v=VIDEOID 59 60 # Download a live stream as it happens 61 yt-dlp --live-from-start https://www.youtube.com/watch?v=LIVEID 62 ``` 63 64 > [!tip]+ Bot Detection Bypass Methods 65 > `fas:Lightbulb` 66 > When YouTube shows "Sign in to confirm you're not a bot": 67 > ```bash 68 > # Method 1: Use cookies from your logged-in browser 69 > yt-dlp --cookies-from-browser firefox https://www.youtube.com/watch?v=VIDEOID 70 > 71 > # Method 2: Spoof the Android client 72 > yt-dlp --extractor-args "youtube:player_client=android" https://www.youtube.com/watch?v=VIDEOID 73 > 74 > # Method 3: Combine both (most reliable) 75 > yt-dlp --cookies-from-browser firefox --extractor-args "youtube:player_client=android,web" https://www.youtube.com/watch?v=VIDEOID 76 > ``` 77 78 **Other platform downloads:** 79 80 ```bash 81 # Twitter/X videos 82 yt-dlp https://x.com/user/status/TWEETID 83 84 # Twitch VOD 85 yt-dlp https://www.twitch.tv/videos/VODID 86 87 # Twitch clip 88 yt-dlp https://clips.twitch.tv/CLIPNAME 89 90 # TikTok individual video 91 yt-dlp https://www.tiktok.com/@user/video/VIDEOID 92 93 # Instagram post (requires cookies) 94 yt-dlp --cookies-from-browser firefox https://www.instagram.com/p/POSTID/ 95 96 # Facebook video (requires cookies) 97 yt-dlp --cookies-from-browser firefox https://www.facebook.com/video/VIDEOID 98 99 # Twitter Space → optimised MP3 100 yt-dlp -x --audio-format mp3 --audio-quality 64K https://twitter.com/i/spaces/SPACEID 101 102 # Spotify podcast episode 103 yt-dlp https://open.spotify.com/episode/EPISODEID 104 ``` 105 106 --- 107 108 ### ffmpeg — Video Processing `fas:Terminal` 109 110 > [!info]+ [ffmpeg](https://ffmpeg.org) Overview 111 > `fas:Terminal` 112 > Essential standalone tool for re-encoding, compressing, splitting, and converting video files. Required by yt-dlp. 113 > 1. Converts any format to universally compatible H.264 MP4 114 > 2. GPU-accelerated encoding via NVENC (NVIDIA) 115 > 3. Lossless stream copy cuts and metadata embedding 116 117 **Universal re-encode to H.264 MP4 (works on ANY input format):** 118 119 ```bash 120 ffmpeg -hwaccel auto -i YOUR_INPUT_FILE.anything \ 121 -c:v libx264 \ 122 -pix_fmt yuv420p \ 123 -profile:v baseline \ 124 -level 3.0 \ 125 -crf 22 \ 126 -preset medium \ 127 -c:a aac \ 128 -strict experimental \ 129 -movflags +faststart \ 130 -threads 0 \ 131 output.mp4 132 ``` 133 134 > [!info]+ CRF Quality Guide 135 > 1. **CRF 18** = near-lossless — use for archival masters 136 > 2. **CRF 22** = good quality — default for most archiving 137 > 3. **CRF 28** = smaller file, visible quality loss — throwaway clips only 138 > 4. *Lower number = better quality + larger file size* 139 140 **With downscaling to 720p (reduces file size significantly):** 141 142 ```bash 143 ffmpeg -hwaccel auto -i input.mp4 \ 144 -c:v libx264 -pix_fmt yuv420p -profile:v baseline -level 3.0 \ 145 -crf 22 -vf scale=-2:720 -preset medium \ 146 -c:a aac -strict experimental -movflags +faststart -threads 0 \ 147 output.mp4 148 ``` 149 150 *Replace `720` with `1080`, `480`, or `360` as needed. The `-2` ensures width is divisible by 2 (required for MP4).* 151 152 **GPU-accelerated encoding (NVIDIA NVENC):** 153 154 ```bash 155 ffmpeg -hwaccel cuda -i input.mp4 -c:v h264_nvenc -cq 22 -c:a aac output.mp4 156 # Note: use -cq (not -crf) for NVENC constant quality mode 157 ``` 158 159 **Split large video into 1-hour chunks:** 160 161 ```bash 162 ffmpeg -hwaccel auto -i input.mp4 \ 163 -c copy -map 0 \ 164 -segment_time 01:00:00 \ 165 -f segment \ 166 -reset_timestamps 1 \ 167 -movflags +faststart \ 168 -threads 0 \ 169 output%03d.mp4 170 # Produces output000.mp4, output001.mp4, etc. 171 ``` 172 173 **Quick lossless cut (no re-encode, very fast):** 174 175 ```bash 176 ffmpeg -i input.mp4 -ss 00:01:00 -to 00:05:00 -c copy output.mp4 177 ``` 178 179 **Embed metadata/chapters into a downloaded video:** 180 181 ```bash 182 ffmpeg -i input.mp4 -i metadata.txt -map_metadata 1 -c copy output.mp4 183 ``` 184 185 > [!tip]+ Video Size Strategy 186 > `fas:Lightbulb` 187 > 1. **Under 200MB** → upload directly, no processing needed 188 > 2. **200MB–400MB** → re-encode at 720p first, usually gets under 200MB 189 > 3. **400MB+** → re-encode first, then split if still too large 190 > 4. **Never split without re-encoding first** — don't upload large unoptimised chunks 191 > 5. **Audio quality matters more than video** — reduce resolution before touching audio bitrate 192 193 --- 194 195 ### Online Video Archiving (No Command Line) 196 197 | Tool | URL | Notes | 198 |---|---|---| 199 | **PreserveTube** | [preservetube.com](https://www.preservetube.com) | YouTube up to ~2 hours. Has a UserScript for a YouTube button. | 200 | **Cobalt** | [cobalt.tools](https://cobalt.tools) | YouTube, Twitter, TikTok, Instagram. ~100 min limit. | 201 | **Twitter Video DL** | [twittervideodownloader.com](https://twittervideodownloader.com) | Simple Twitter/X downloads, no install. | 202 | **YouTube Multi DL** | [youtubemultidownloader.net](https://youtubemultidownloader.net) | Bulk download multiple YouTube videos. | 203 | **Filmot** | [filmot.com](https://filmot.com) | **Finds unlisted and deleted YouTube videos** by searching auto-captions. | 204 | **Tartube** | [github.com/axcore/tartube](https://github.com/axcore/tartube) | GUI wrapper for yt-dlp — cross-platform. | 205 | **Handbrake** | [handbrake.fr](https://handbrake.fr) | GUI video compressor (ffmpeg/x264 wrapper). RF 20–22 recommended. | 206 207 --- 208 209 ## Part 3 — Screenshots & Full-Page Captures 210 211 ### Platform Quick Reference 212 213 | Platform | Method | Shortcut | 214 |---|---|---| 215 | Windows | Snipping Tool | `Win + Shift + S` | 216 | Windows (advanced) | ShareX | [getsharex.com](https://getsharex.com) — blur, annotate, upload | 217 | macOS | Full screen | `Cmd + Shift + 3` | 218 | macOS | Select area | `Cmd + Shift + 4` | 219 | macOS | Specific window | `Cmd + Shift + 4` then `Space` | 220 | iPhone (old) | Power + Home | Saves to Camera Roll | 221 | iPhone (new) | Power + Vol Up | Saves to Camera Roll | 222 | Firefox | Built-in | Right-click → Take Screenshot → Save Full Page | 223 | Chrome/Brave | DevTools | `F12` → `Ctrl+Shift+P` → "Capture full size screenshot" | 224 225 > [!tip]+ Firefox Developer Console Screenshots 226 > `fas:Terminal` 227 > Press `Shift + F2` to open the developer toolbar, then: 228 > ``` 229 > screenshot --fullpage # saves to Downloads 230 > screenshot --fullpage --clipboard # copies to clipboard 231 > screenshot --fullpage myfilename # saves with custom name 232 > screenshot --fullpage --delay 3 # 3 second delay before capture 233 > ``` 234 235 > [!tip]+ PNG Compression — Reduce File Size Before Uploading 236 > `fas:Lightbulb` 237 > 1. **pngquant** (CLI): `pngquant --quality=65-80 screenshot.png` 238 > 2. **Squoosh** (browser): [squoosh.app](https://squoosh.app) 239 > 3. **TinyPNG** (browser): [tinypng.com](https://tinypng.com) 240 241 --- 242 243 ## Part 4 — Platform-Specific Archiving 244 245 ### Twitter / X 246 247 > [!warning]+ Twitter/X requires login for most content — breaks most archivers 248 > Methods in order of reliability: 249 > 1. **Nitter mirror → archive.today**: Replace `twitter.com`/`x.com` with `nitter.poast.org`, feed to archive.ph 250 > 2. **GhostArchive** directly on x.com URL — works intermittently 251 > 3. **Thread Reader** for multi-tweet threads: `https://threadreaderapp.com/thread/TWEETID` → archive the Thread Reader URL on archive.ph (static HTML, archives perfectly) 252 > 4. **yt-dlp** for videos: `yt-dlp https://x.com/user/status/TWEETID` 253 > 5. **Twint** for bulk account archiving (partially broken): [github.com/twintproject/twint](https://github.com/twintproject/twint) 254 > 6. **Megalodon** — currently best for archiving individual X/Twitter page URLs 255 256 > [!danger]+ Facebook Image Download URLs Contain Your Identity Token 257 > `fas:Skull` 258 > Facebook image download URLs contain an identifying token that can be traced back to your account. Always rename the file or copy-paste the image rather than downloading directly. 259 260 --- 261 262 ### Reddit 263 264 > [!info]+ Reddit Archiving 265 > 1. Always use **old.reddit.com** URLs — `reddit.com` → `old.reddit.com`. archive.today handles old Reddit layout far better. 266 > 2. **Unddit** for deleted posts/comments: replace `reddit.com` with `unddit.com` in any URL → [unddit.com](https://unddit.com) 267 > 3. **Full user post history**: use PRAW (Python Reddit API Wrapper) to iterate profile pages and submit each to archive.today. *See `scripts/reddit_archive_user.py`.* 268 269 --- 270 271 ### Discord 272 273 > [!info]+ [DiscordChatExporter](https://github.com/Tyrrrz/DiscordChatExporter) — Standard Discord archiving tool 274 > Exports Discord servers, channels, and DMs to HTML (with images), JSON, CSV, or TXT. GUI + CLI versions. 275 > 1. Requires your Discord auth token (browser DevTools → Network tab → Authorization header) 276 > 2. GUI version is straightforward; CLI supports bulk and automated export 277 278 ```bash 279 # CLI bulk export of an entire server 280 DiscordChatExporter.Cli exportguild --guild SERVERID --token YOUR_TOKEN --output ./export/ --format HtmlDark 281 ``` 282 283 > [!warning]+ Discord CDN URLs Expire 284 > `cdn.discordapp.com` image URLs are publicly accessible without login, but Discord periodically rotates/expires them. Archive immediately after finding them. 285 286 --- 287 288 ### Instagram 289 290 > [!info]+ Instagram Archiving Methods 291 > Instagram requires login for most content. Working methods: 292 > 1. **View profiles without account**: [picuki.com](https://picuki.com), [imgsed.com](https://imgsed.com), [dumpor.com](https://dumpor.com) → then feed URL to archive.today 293 > 2. **Download posts/reels**: [Instaloader](https://instaloader.github.io) (`instaloader profile USERNAME`), JDownloader, [SnapInsta](https://snapinsta.app), [igram.io](https://igram.io) 294 > 3. **Stories** (disappear after 24h — archive immediately): Instaloader with `--stories` flag (requires login) 295 > 4. **With yt-dlp** (requires cookies): `yt-dlp --cookies-from-browser firefox https://www.instagram.com/p/POSTID/` 296 297 > [!danger]+ Instagram Screenshot Self-Doxx Warning 298 > `fas:Skull` 299 > Instagram embeds your avatar and username into the comment field UI. **Crop or redact your username from every Instagram screenshot before posting.** 300 301 --- 302 303 ### TikTok 304 305 ```bash 306 # Individual video (works) 307 yt-dlp https://www.tiktok.com/@user/video/VIDEOID 308 ``` 309 310 > [!warning]+ yt-dlp profile scraping is broken for TikTok 311 > Individual videos still work. For profile-level scraping, use [Geranium's scraper-helper](https://github.com/GeraniumKF/scraper-helper) or [Cobalt](https://cobalt.tools). 312 313 --- 314 315 ### YouTube (Special Cases) 316 317 ```bash 318 # Age-restricted (requires logged-in cookies) 319 yt-dlp --cookies-from-browser chrome https://www.youtube.com/watch?v=VIDEOID 320 321 # Full channel backup with metadata + organised directory structure 322 yt-dlp --ignore-errors --continue \ 323 --write-description --write-info-json \ 324 --write-thumbnail --write-subs \ 325 --output "%(uploader)s/%(upload_date)s-%(title)s.%(ext)s" \ 326 https://www.youtube.com/c/CHANNELNAME 327 ``` 328 329 > [!tip]+ Finding Deleted/Unlisted YouTube Videos 330 > `fas:Lightbulb` 331 > [Filmot](https://filmot.com) indexes YouTube subtitle data including from removed videos. Search by keywords to find videos no longer publicly visible — extremely useful for tracking deleted content. 332 333 > [!info]+ YouTube Comment Archiving 334 > No third-party site currently preserves full comment sections. Use the [YouTube Data API v3](https://developers.google.com/youtube/v3) (free API key) to download all comments from a video programmatically. *See `scripts/youtube_comment_archiver.py`.* 335 336 --- 337 338 ### Other Platforms 339 340 > [!info]+ Platform Reference Table 341 > 342 > | Platform | Tool / Method | 343 > |---|---| 344 > | **Twitch VODs** | `yt-dlp https://www.twitch.tv/videos/VODID` | 345 > | **Twitch Clips** | `yt-dlp https://clips.twitch.tv/CLIPNAME` | 346 > | **Twitch Live** | `yt-dlp https://www.twitch.tv/CHANNELNAME` | 347 > | **Steam** | [steamid.io](https://steamid.io) to get numeric ID; archive both vanity + `/profiles/76561...` URLs | 348 > | **DeviantArt** | [RipMe](https://github.com/RipMeApp/ripme) (requires Java JDK) — archives entire galleries | 349 > | **Tumblr** | archive.today handles NSFW blogs; original-posts-only: [studiomoh.com/fun/tumblr_originals](https://studiomoh.com/fun/tumblr_originals) | 350 > | **4chan** | [4plebs.org](https://4plebs.org) — permanent archive. Threads 404 fast — archive immediately. Also archive the 4plebs link itself on archive.today. | 351 > | **Gab** | [garc](https://github.com/DocNow/garc) — Gab API scraper (fork of twarc), requires Gab login | 352 > | **Bluesky** | archive.today works well with Bluesky URLs currently | 353 > | **AO3 / FFN** | FanFictionDownloader — downloads with full metadata (URL, timestamp, author) | 354 > | **Spotify / Podcasts** | `yt-dlp https://open.spotify.com/episode/EPISODEID` (some require logged-in session) | 355 > | **Facebook Videos** | `yt-dlp --cookies-from-browser firefox https://www.facebook.com/video/VIDEOID` (click through content warning first) | 356 357 --- 358 359 ## Part 5 — Bulk & Automated Archiving `fas:ClipboardList` 360 361 > [!info]+ Batch Submit URLs to Wayback Machine 362 > `fas:Terminal` 363 > ```bash 364 > # Submit a list of URLs from a file 365 > cat urls.txt | xargs -I{} curl -s "https://web.archive.org/save/{}" 366 > ``` 367 > *See `scripts/batch_archive.sh` for a rate-limited, logged version of this.* 368 369 > [!info]+ [Internet Archive CLI](https://github.com/jjjake/internetarchive) 370 > `fas:Terminal` 371 > Official CLI for uploading files directly to the Internet Archive. 372 > ```bash 373 > pip install internetarchive 374 > ia upload IDENTIFIER /path/to/file 375 > ``` 376 377 > [!info]+ [Bellingcat Auto Archiver](https://github.com/bellingcat/auto-archiver) 378 > Professional-grade automated archiving. Takes URLs from spreadsheets, Telegram, or other sources and archives to archive.org, Google Drive, S3, etc. with verification metadata. Used by journalists and OSINT researchers. 379 380 > [!info]+ Archiving Entire Wikis and Forums 381 > `fas:Terminal` 382 > ```bash 383 > # HTTrack (recommended — handles link rewriting automatically) 384 > httrack https://wiki.example.com -O ./local_mirror/ -r6 385 > 386 > # wget (alternative) 387 > wget --mirror --convert-links --adjust-extension --page-requisites --no-parent https://wiki.example.com 388 > ``` 389 > For large-scale professional crawls: [Heritrix](https://github.com/internetarchive/heritrix3) — the Internet Archive's own open-source crawler. 390 391 > [!info]+ Download an Existing Wayback Machine Snapshot Locally 392 > `fas:Terminal` 393 > ```bash 394 > gem install wayback_machine_downloader 395 > wayback_machine_downloader http://example.com 396 > 397 > # Specify a date range 398 > wayback_machine_downloader http://example.com --from 20200101 --to 20201231 399 > ``` 400 401 --- 402 403 ## Part 6 — OSINT & Account Finding 404 405 | Tool | URL | Use For | 406 |---|---|---| 407 | **Sherlock** | [github.com/sherlock-project/sherlock](https://github.com/sherlock-project/sherlock) | Username search across 300+ social media sites | 408 | **Maigret** | [github.com/soxoj/maigret](https://github.com/soxoj/maigret) | More comprehensive than Sherlock, shows profile info | 409 | **Filmot** | [filmot.com](https://filmot.com) | Find unlisted/deleted YouTube videos via subtitle search | 410 | **vacbanned.com** | [vacbanned.com](http://www.vacbanned.com) | Names associated with VAC-banned Steam accounts | 411 | **SteamID.io** | [steamid.io](https://steamid.io) | Convert Steam vanity URLs to permanent numeric IDs | 412 413 ```bash 414 # Sherlock 415 python3 sherlock username 416 417 # Maigret (more detailed, more sites) 418 python3 -m maigret username 419 ``` 420 421 --- 422 423 ## Part 7 — File Hosting for Large Files 424 425 > [!tip]+ File Hosting Priority Order 426 > `fas:Lightbulb` 427 > 1. **On-site upload** — always preferred. Use your platform's native upload (e.g. 200MB limit on most forums). 428 > 2. **MEGA** ([mega.nz](https://mega.nz)) — widely accepted off-site host. Large files, free tier. 429 > 3. **Internet Archive** ([archive.org/upload](https://archive.org/upload)) — free, unlimited, permanent. Best for large video files. 430 > 4. **multiup.io** ([multiup.io](https://multiup.io)) — upload once, mirrors automatically to multiple file hosts. Redundant links. 431 > 432 > **Avoid:** Any single-purpose file hosts, social media embeds, or link shorteners. They disappear. 433 434 --- 435 436 ## Part 8 — Un-Redacting & Recovering Obscured Content 437 438 ### Recovering Poorly Redacted Text 439 440 > [!tip]+ Text Hidden Under a Black Bar (Image Editor Method) 441 > `fas:Lightbulb` 442 > If text was redacted with a black rectangle placed *over* it (not burned in): 443 > 1. Open in GIMP, Photoshop, or even MS Paint 444 > 2. Try **Levels** or **Curves** adjustment — drag input levels to extremes 445 > 3. Try **Contrast** at maximum 446 > 4. If the black bar is a separate layer (common in quick edits), select and delete the layer 447 448 ### Recovering Deleted Content 449 450 | Method | Where to Look | 451 |---|---| 452 | **Google Cache** | `cache:https://example.com/page` (being phased out, sometimes works) | 453 | **Wayback Machine** | `https://web.archive.org/web/*/example.com/page` | 454 | **Unddit** (Reddit) | Replace `reddit.com` with `unddit.com` | 455 | **Filmot** (YouTube) | Search deleted video titles/content at [filmot.com](https://filmot.com) | 456 | **Google snippets** | Search for the title — results show text excerpts even when cache is gone | 457 458 --- 459 460 ## Part 9 — Expanded: Cookie Extraction for Archiving 461 462 Many archiving tasks require authenticated sessions. The cleanest method is exporting cookies from your browser. 463 464 > [!info]+ Exporting Cookies with yt-dlp (Automatic) 465 > `fas:Terminal` 466 > yt-dlp can pull cookies directly from your browser — no manual export needed: 467 > ```bash 468 > yt-dlp --cookies-from-browser firefox URL 469 > yt-dlp --cookies-from-browser chrome URL 470 > yt-dlp --cookies-from-browser brave URL 471 > ``` 472 473 > [!info]+ Manual Cookie Export (Netscape Format) 474 > `fas:Terminal` 475 > For tools that require a `cookies.txt` file: 476 > 1. Install the **Get cookies.txt LOCALLY** extension in Firefox or Chrome 477 > 2. Navigate to the site while logged in 478 > 3. Click the extension → export → save as `cookies.txt` 479 > 4. Use with: `yt-dlp --cookies cookies.txt URL` 480 481 > [!warning]+ Cookie Security 482 > 1. Never share your cookies.txt file — it grants full session access to your accounts 483 > 2. Delete exported cookie files after use 484 > 3. Use throwaway accounts for archiving sensitive content 485 486 --- 487 488 ## Part 10 — Expanded: VTT Subtitle Processing 489 490 Auto-generated subtitles from yt-dlp are invaluable for searching long videos and streams. 491 492 ```bash 493 # Download auto-generated subtitles alongside video 494 yt-dlp --write-auto-subs --sub-format vtt https://www.youtube.com/watch?v=VIDEOID 495 496 # Download subtitles only (no video) 497 yt-dlp --skip-download --write-auto-subs --sub-format vtt https://www.youtube.com/watch?v=VIDEOID 498 ``` 499 500 *See `scripts/vtt_to_text.py` for a script that converts `.vtt` files to searchable plaintext.* 501 502 --- 503 504 ## Complete Tool Reference `fas:ClipboardList` 505 506 ### Web Archiving 507 508 | Tool | URL | Best For | 509 |---|---|---| 510 | archive.today | [archive.ph](https://archive.ph) | Primary web page archiving | 511 | GhostArchive | [ghostarchive.org](https://ghostarchive.org) | Backup archiver, LinkedIn, Facebook | 512 | Megalodon | [megalodon.jp](https://megalodon.jp) | Japanese archiver, good Twitter backup | 513 | Wayback Machine | [web.archive.org](https://web.archive.org) | Finding old archives — not primary | 514 | Perma.cc | [perma.cc](https://perma.cc) | Legal-grade citations | 515 | archiveweb.page | [archiveweb.page](https://archiveweb.page) | Local WARC archives of JS-heavy sites | 516 | Browsertrix | [browsertrix.com](https://browsertrix.com) | Automated high-fidelity crawling | 517 | monolith | [github.com/Y2Z/monolith](https://github.com/Y2Z/monolith) | Single-file local HTML archive | 518 | HTTrack | [httrack.com](http://httrack.com) | Full site mirror with link rewriting | 519 520 ### Video Downloading 521 522 | Tool | URL | Best For | 523 |---|---|---| 524 | yt-dlp | [github.com/yt-dlp/yt-dlp](https://github.com/yt-dlp/yt-dlp) | Everything | 525 | ffmpeg | [ffmpeg.org](https://ffmpeg.org) | Re-encoding, splitting, converting | 526 | PreserveTube | [preservetube.com](https://www.preservetube.com) | Easy YouTube archiving online | 527 | Cobalt | [cobalt.tools](https://cobalt.tools) | Quick multi-platform downloads | 528 | Tartube | [github.com/axcore/tartube](https://github.com/axcore/tartube) | yt-dlp GUI | 529 | Handbrake | [handbrake.fr](https://handbrake.fr) | GUI video compression | 530 | Filmot | [filmot.com](https://filmot.com) | Find deleted/unlisted YouTube videos | 531 | Twitter Video DL | [twittervideodownloader.com](https://twittervideodownloader.com) | Twitter/X videos without tools | 532 533 ### OSINT 534 535 | Tool | URL | Best For | 536 |---|---|---| 537 | Sherlock | [github.com/sherlock-project/sherlock](https://github.com/sherlock-project/sherlock) | Username enumeration across 300+ sites | 538 | Maigret | [github.com/soxoj/maigret](https://github.com/soxoj/maigret) | Comprehensive username + profile OSINT | 539 | Filmot | [filmot.com](https://filmot.com) | Deleted YouTube video discovery | 540 | SteamID.io | [steamid.io](https://steamid.io) | Steam vanity URL → permanent ID | 541 542 --- 543 544 ## Lessons Learned `fas:Lightbulb` 545 546 1. **Redundancy is everything.** A single archiver is a single point of failure. Submit every important page to at least archive.today AND one backup (GhostArchive or Megalodon). Services die, get pressured, or get acquired. 547 2. **Old interfaces archive better.** `old.reddit.com` vs new Reddit, Nitter vs Twitter.com — JavaScript-heavy modern UIs break archivers. Always prefer the legacy URL when one exists. 548 3. **Cookies solve most login walls.** yt-dlp's `--cookies-from-browser` flag handles age restrictions, rate limits, and member-only content without manual token extraction. 549 4. **Audio > video when cutting file size.** When you need to reduce a video for upload, drop resolution (1080p → 720p) before touching audio bitrate. Compressed audio is noticeably worse; compressed video is not. 550 5. **Archive the archive.** External archive links (4plebs, Wayback, etc.) can themselves go down. If something is critical, archive the archive URL on archive.today too. 551 552 --- 553 554 ## References `fas:BookOpen` 555 556 1. [yt-dlp GitHub](https://github.com/yt-dlp/yt-dlp) 557 2. [ffmpeg Official Download](https://ffmpeg.org/download.html) 558 3. [archive.today / archive.ph](https://archive.ph) 559 4. [GhostArchive](https://ghostarchive.org) 560 5. [Megalodon](https://megalodon.jp) 561 6. [Wayback Machine](https://web.archive.org) 562 7. [Perma.cc](https://perma.cc) 563 8. [archiveweb.page (Webrecorder)](https://archiveweb.page) 564 9. [Browsertrix Crawler](https://github.com/webrecorder/browsertrix-crawler) 565 10. [monolith](https://github.com/Y2Z/monolith) 566 11. [SingleFile Firefox Extension](https://addons.mozilla.org/en-US/firefox/addon/single-file/) 567 12. [HTTrack](http://httrack.com) 568 13. [DiscordChatExporter](https://github.com/Tyrrrz/DiscordChatExporter) 569 14. [Instaloader](https://instaloader.github.io) 570 15. [Sherlock](https://github.com/sherlock-project/sherlock) 571 16. [Maigret](https://github.com/soxoj/maigret) 572 17. [Filmot — YouTube subtitle search](https://filmot.com) 573 18. [Bellingcat Auto Archiver](https://github.com/bellingcat/auto-archiver) 574 19. [Internet Archive CLI](https://github.com/jjjake/internetarchive) 575 20. [Heritrix Web Crawler](https://github.com/internetarchive/heritrix3) 576 21. [Thread Reader App](https://threadreaderapp.com) 577 22. [Unddit — deleted Reddit posts](https://unddit.com) 578 23. [SteamID.io](https://steamid.io) 579 24. [Cobalt Downloader](https://cobalt.tools) 580 25. [PreserveTube](https://www.preservetube.com) 581 26. [Tartube](https://github.com/axcore/tartube) 582 27. [ShareX](https://getsharex.com) 583 28. [Handbrake](https://handbrake.fr) 584 29. [RipMe](https://github.com/RipMeApp/ripme) 585 30. [Twint](https://github.com/twintproject/twint) 586 31. [YouTube Data API v3](https://developers.google.com/youtube/v3) 587 32. [4plebs](https://4plebs.org) 588 33. [garc — Gab scraper](https://github.com/DocNow/garc) 589 34. [Nitter mirror](https://nitter.poast.org) 590 35. [webvtt-py](https://github.com/glut23/webvtt-py) 591 36. [multiup.io](https://multiup.io) 592 37. [MEGA](https://mega.nz) 593 38. [Squoosh](https://squoosh.app) 594 595 --- 596 597 #Archival #OSINT #Tools #Cheatsheet #yt-dlp #ffmpeg #WebArchiving #VideoDownloading #ContentRecovery