social-media-platforms.md (85401B)
1 --- 2 title: "Social Media — Per-Platform Research" 3 description: "What each major platform still exposes to an unauthenticated researcher, and the tools that work per platform." 4 category: osint 5 subcategory: "Social Media" 6 tags: [osint, social-media, telegram, tiktok] 7 tools: [yt-dlp, gallery-dl, instaloader, telethon, telepathy, arctic-shift, atproto, goat, tiktok-hashtag-analysis, instagram-location-search] 8 difficulty: intermediate 9 updated: 2026-10-04 10 references: 11 - name: "Bellingcat's Online Investigation Toolkit" 12 url: "https://bellingcat.gitbook.io/toolkit" 13 author: "Bellingcat" 14 license: none 15 relation: derived 16 note: "Tool catalogue: names, descriptions, cost flags and links for this area." 17 - name: "OSINT Newsletter Tools Library" 18 url: "https://tools.osintnewsletter.com" 19 author: "The OSINT Newsletter" 20 license: none 21 relation: derived 22 note: "Second tool catalogue, cross-checked against the above." 23 - name: "yt-dlp option reference" 24 url: "https://github.com/yt-dlp/yt-dlp#usage-and-options" 25 author: "yt-dlp contributors" 26 relation: link-only 27 note: "Every yt-dlp flag on this page was checked against the project's own option list." 28 - name: "gallery-dl command-line options" 29 url: "https://github.com/mikf/gallery-dl/blob/master/docs/options.md" 30 author: "Mike Fährmann" 31 relation: link-only 32 note: "Every gallery-dl flag on this page was checked against the project's own option list." 33 - name: "Instaloader command-line options" 34 url: "https://instaloader.github.io/cli-options.html" 35 author: "Alexander Graf, André Koch-Kramer" 36 relation: link-only 37 note: "Instaloader flags and target syntax checked against the project's own documentation." 38 - name: "Arctic Shift API reference" 39 url: "https://github.com/ArthurHeitmann/arctic_shift/blob/master/api/README.md" 40 author: "Arthur Heitmann" 41 relation: link-only 42 note: "Endpoint paths and query parameters checked against the project's own API reference." 43 - name: "AT Protocol specifications and lexicons" 44 url: "https://atproto.com/specs/tid" 45 author: "Bluesky Social" 46 relation: link-only 47 note: "XRPC parameter names and the TID timestamp layout checked against the published lexicons." 48 - name: "Meta Ad Library API reference" 49 url: "https://developers.facebook.com/docs/graph-api/reference/ads_archive/" 50 author: "Meta Platforms" 51 relation: link-only 52 note: "ads_archive parameters, enum values and field names checked against the API reference." 53 --- 54 55 ## What this covers 56 57 Each platform exposes a different surface and closes it on a different schedule. This is the 58 per-platform picture: what is still reachable without an account, what needs one, and which tools 59 survive contact with the current APIs. Assume anything here can break without notice — platforms 60 change access rules faster than tooling keeps up. 61 62 ## Method 63 64 1. **Decide whether you can be seen before you touch anything.** Some platforms attribute a view, 65 a follow or a group join to your account. Work out which, on this platform, today — then pick 66 the account you are willing to burn. 67 2. **Archive first, analyse second.** Capture the post, the profile and the surrounding thread 68 before you start pulling threads. Accounts get deleted mid-investigation, routinely. 69 3. **Take the public surface without an account** where it exists: profile metadata, follower 70 counts, pinned posts, video metadata. Much of this is still reachable unauthenticated and 71 costs you nothing. 72 4. **Pivot on identifiers, not display names.** Numeric IDs, DIDs and channel IDs survive renames 73 and vanity-URL changes; handles and names do not. 74 5. **Trace to origin.** The account with the most engagement is almost never the source. Follow 75 forwards, quote-posts and reuploads backwards until the chain stops. 76 6. **Date everything from the platform, not the page.** Upload timestamps, message IDs and 77 creation dates are platform-side facts; the date shown on a reposted screenshot is not. 78 7. **Expect the tooling to be broken.** Anything built on an undocumented endpoint dies without 79 notice. Check the project's last commit before you trust its output, and cross-check one result 80 by hand. 81 82 The judgement calls: decide early whether you need the content or the network, because bulk media 83 capture and graph enumeration use different tools and different amounts of patience. Decide which 84 account you are prepared to attribute the work to, before you authenticate anything. And decide 85 whether a date is given or derived — if you read a date off the page and then "confirm" it with a 86 tool that reads the same page, you have proved nothing. 87 88 ## The identifier arithmetic 89 90 Most platform IDs are not counters. They are clock readings, so the ID you already have dates the 91 object with no request to the platform at all. That matters twice over: the object may be deleted 92 by the time you look, and the date rendered on a page is frequently the repost's date rather than 93 the original's. 94 95 **TikTok.** A video ID is a 64-bit value whose top 32 bits hold a Unix timestamp in seconds. Shift 96 right by 32 and read it: 97 98 ```text 99 video id : 7169068344567909638 100 width : 63 bits, so the timestamp occupies bits 62..32 101 id >> 32 : 1669178797 102 as UTC : 2022-11-23 04:46:37Z 103 ``` 104 105 **X / Twitter.** Status IDs are Snowflakes: 41 bits of milliseconds since a platform-specific 106 epoch, sitting 22 bits up. The epoch is the number you have to get right. 107 108 ```text 109 status id : 1949394974262350098 110 id >> 22 : 464771979871 milliseconds since the Twitter epoch 111 + 1288834974657 : 1753606954528 milliseconds since the Unix epoch 112 as UTC : 2025-07-27 09:02:34.528Z 113 ``` 114 115 **Bluesky.** A record key is a TID: 13 characters of the sortable base32 alphabet 116 `234567abcdefghijklmnopqrstuvwxyz`, decoding to a top bit of 0, then 53 bits of **microseconds** 117 since the Unix epoch, then a 10-bit random clock identifier. Because the alphabet sorts, record 118 keys sort chronologically as plain strings. 119 120 ```bash 121 # all three, from the shell 122 python3 -c 'import datetime as d; i=7169068344567909638; print(d.datetime.fromtimestamp(i>>32, d.timezone.utc))' 123 python3 -c 'import datetime as d; i=1949394974262350098; print(d.datetime.fromtimestamp(((i>>22)+1288834974657)/1000, d.timezone.utc))' 124 goat syntax tid inspect 3kzifvcppte22 125 ``` 126 127 ```text 128 2022-11-23 04:46:37+00:00 129 2025-07-27 09:02:34.528000+00:00 130 Timestamp (UTC): 2024-08-12T02:08:03.29Z 131 Timestamp (Local): 2024-08-11T19:08:03-07:00 132 ClockID: 0 133 uint64: 0x187dcbda2b5ca800 134 ``` 135 136 Three ways this misleads. The first is the obvious one: an ID dates the **record**, not the 137 footage. A clip shot in 2019 and uploaded in 2026 has a 2026 ID, and the arithmetic is still 138 correct. The second is who generated the number. TikTok and X mint IDs server-side at creation, so 139 they are harder to forge than a rendered date; an atproto TID is minted by whichever component 140 writes the record, off that generator's own local clock, and the specification states plainly that 141 uniqueness cannot be guaranteed because an adversarial party can create records reusing known TIDs. 142 A Bluesky record key is therefore the same class of evidence as `record.createdAt` — an assertion 143 by the writer — and the relay's `indexedAt` is the independent number. The third is bit-width drift: quote the shift, not a 144 string slice. Taking "the first 31 characters of the binary" works for every current 19-digit 145 TikTok ID only because such IDs happen to be 63 bits wide; `id >> 32` keeps working when they are 146 not. 147 148 ## Telegram 149 150 The most productive platform for open-source research, because public channels are genuinely 151 public, history is retained, and forwarding metadata exposes the network between channels. 152 153 - **Forward chains** are the key primitive. A forwarded message keeps its origin, so you can trace 154 content back through the channels that amplified it. 155 - **Member lists** are visible in many groups. Your own account appears in them too — use a 156 research account. 157 - **Channel creation dates and message IDs** are sequential, which lets you estimate when a channel 158 started and how much has been deleted. 159 160 Before any tooling, the web preview gives you a channel's recent history with no account, no app 161 and no API credentials. It is the cheapest first look there is, and it is the only one that costs 162 you nothing in attribution: 163 164 ```text 165 https://t.me/s/channelname the public web view: recent messages, dates, view counts 166 https://t.me/channelname/4312 one message, by its sequential id -- the permalink form 167 https://t.me/s/channelname?before=4312 page backwards from that id, ~20 messages at a time 168 https://t.me/+AbCdEf0123456789 a private invite link; opening it does not join, but 169 it does reveal the chat title and member count 170 ``` 171 172 ```bash 173 # how far back the web view goes, and whether ids are contiguous 174 curl -s 'https://t.me/s/channelname' \ 175 | grep -o 'data-post="channelname/[0-9]*"' | sed 's/.*\///; s/"//' | sort -n | head -1 176 177 # walk backwards in pages of ~20 and collect every id the preview will serve 178 for b in 4312 4292 4272 4252; do 179 curl -s "https://t.me/s/channelname?before=$b" \ 180 | grep -o 'data-post="channelname/[0-9]*"' | sed 's/.*\///; s/"//' 181 done | sort -nu > ids.txt 182 183 # the gaps are deletions: compare the id range against the count you actually got 184 awk 'NR==1{min=$1} {max=$1; n++} END{print "ids", min"-"max, "present", n, "missing", max-min+1-n}' ids.txt 185 ``` 186 187 ```text 188 ids 401-440 present 39 missing 1 189 ``` 190 191 Strip the id with `sed`, not with a second `grep -o '[0-9]*$'`. The attribute ends in a quote, so 192 anchoring digits to end-of-line matches the empty string and emits one blank line per post: twenty 193 matches in, twenty blanks out, exit status 0. The run looks like a channel with no posts rather 194 than like a pipeline that threw the answer away. 195 196 The web view is a *preview*, not the archive: it serves a bounded window of recent history and 197 will not page back indefinitely, and a channel can disable it entirely. Treat a missing ID as 198 evidence of deletion only after you have confirmed the preview is serving that range at all. 199 Archiving a channel properly, mapping its forwards and checking a number against the platform all 200 have tooling below under [Key tools](#key-tools). 201 202 ## Facebook 203 204 Graph search is gone and most enumeration is dead. What still works: 205 206 - **Numeric profile IDs** persist even when vanity URLs change, and map back to a profile. 207 - **Public pages, groups and events** remain readable and are frequently forgotten by their owners. 208 - **Ad Library** ([facebook.com/ads/library](https://www.facebook.com/ads/library/)) is genuinely 209 open, keyword-searchable, and covers political advertising with spend and reach data. 210 - **Photo comments and reactions** on public posts still expose participant lists. 211 212 The two URL forms worth memorising. The first resolves a numeric ID that survived a rename; the 213 second is the Ad Library filter set, which the web interface writes into its own query string, so 214 you can build a link rather than click through four dropdowns: 215 216 ```text 217 # a numeric id back to a profile, which survives every vanity-URL change 218 https://www.facebook.com/profile.php?id=100012345678901 219 220 # the Ad Library UI's own filter parameters, as it constructs them 221 https://www.facebook.com/ads/library/?active_status=all&ad_type=political_and_issue_ads 222 &country=GB&q=climate&search_type=keyword_unordered&media_type=all 223 224 # every ad from one page, which is the attribution-relevant view 225 https://www.facebook.com/ads/library/?active_status=all&ad_type=all&country=GB 226 &view_all_page_id=123456789 227 ``` 228 229 Those are interface parameters observed from the Ad Library itself, not a documented API; the 230 documented parameter set is the Graph endpoint under 231 [Meta Ad Library API](#meta-ad-library-api) below, and that is what to script against. Everything 232 else on Facebook that once had a query form now does not: [Who posted what?](https://whopostedwhat.com/) 233 and [Graph Tips](https://graph.tips/) build search URLs for you, and whether they work on any given 234 day is a question you answer by trying. 235 236 ## Instagram 237 238 - Profile pictures, bios and post counts are visible without login; the feed usually is not. 239 - `?__a=1` style JSON endpoints have been closed off repeatedly — assume they do not work. 240 - [Instaloader](https://instaloader.github.io/) still handles public profile and hashtag 241 downloads, including comments and geotags where present; it has its own section below. 242 - Aggressive use gets the account and the IP rate-limited quickly, then blocked outright. 243 244 Instaloader's target syntax is the part worth learning, because it is how you address something 245 other than a profile. All four forms are documented targets, not flags: 246 247 ```bash 248 profile_name a public profile, or a private one with --login 249 "#hashtag" posts carrying a hashtag; the quotes are usually necessary 250 "%212998617" posts tagged with a numeric location id 251 "@profile_name" the profiles that profile_name follows, i.e. its followees 252 ``` 253 254 Location IDs are the non-obvious pivot: they are numeric, they are stable, and 255 [instagram-location-search](#instagram-location-search) below turns a coordinate into a list of 256 them, which turns "somewhere near this junction" into a set of addressable targets. 257 258 ## X / Twitter 259 260 Heavily restricted since the API changes. Advanced search still works while logged in, and is 261 still the best tool on the platform — the operator set is below under 262 [Key tools](#key-tools). 263 264 Archived snapshots are now often more reliable than the live site for deleted content — go to the 265 [Wayback Machine](https://web.archive.org/) first. 266 267 One thing still works with no session at all: a status ID dates itself. If you have a screenshot 268 with a URL in it, you have the post's creation time whether or not the post survives: 269 270 ```bash 271 # status id -> UTC creation time, from the Snowflake layout above 272 python3 - <<'PY' 273 import datetime 274 for sid in (1949394974262350098, 1234567890123456789): 275 ms = (sid >> 22) + 1288834974657 276 print(sid, datetime.datetime.fromtimestamp(ms / 1000, datetime.timezone.utc).isoformat()) 277 PY 278 ``` 279 280 ```text 281 1949394974262350098 2025-07-27T09:02:34.528000+00:00 282 1234567890123456789 2020-03-02T19:54:56.824000+00:00 283 ``` 284 285 That is a creation time for the ID, which for an original post is the post's time. For a quote-post 286 or a reply it is that reply's time, not the thing it replies to, and the `conversation_id` operator 287 below is how you get back to the root. 288 289 ## TikTok 290 291 - Profiles and individual videos are viewable without an account. 292 - [Bellingcat's TikTok Hashtag Analysis](https://github.com/bellingcat/tiktok-hashtag-analysis) 293 collects posts by hashtag and charts co-occurrence, which is the useful bulk primitive. 294 - Video download and metadata extraction both work through `yt-dlp`. 295 296 Dating a TikTok needs neither an account nor a live video, because the ID carries the time. This is 297 the whole of what [Bellingcat's extractor](https://bellingcat.github.io/tiktok-timestamp) does, and 298 doing it locally means it also works for a URL you only have as text: 299 300 ```bash 301 # pull the id out of the URL and decode it -- no network request at all 302 url='https://www.tiktok.com/@user/video/7169068344567909638' 303 python3 -c " 304 import datetime, re, sys 305 vid = int(re.search(r'/video/(\d+)', sys.argv[1]).group(1)) 306 print(vid, datetime.datetime.fromtimestamp(vid >> 32, datetime.timezone.utc).isoformat()) 307 " "$url" 308 309 # cross-check against what the platform says, which needs the video to still exist 310 yt-dlp --dump-json "$url" | jq '{id,timestamp,upload_date,uploader_id,view_count}' 311 ``` 312 313 ```text 314 7169068344567909638 2022-11-23T04:46:37+00:00 315 ``` 316 317 When the decoded time and `yt-dlp`'s `timestamp` disagree by more than a second or two, the 318 interesting case is a reupload: the ID in the URL belongs to the copy you are looking at, and an 319 earlier copy with an earlier ID is what you are actually hunting for. 320 321 ## YouTube 322 323 Still the most researcher-friendly large platform: metadata, channel listings, comments and 324 auto-generated subtitles all come out of `yt-dlp`, covered below. 325 326 Upload timestamps are in the metadata and are a reliable earliest-possible date for footage. 327 328 The two invocations that do most of the work on this platform are a channel listing and a subtitle 329 dump — one tells you what exists, the other turns hours of video into text you can grep: 330 331 ```bash 332 # the whole channel as a dated listing, without touching a single video file 333 yt-dlp --flat-playlist --dump-json 'https://youtube.com/@channel/videos' \ 334 | jq -r '[.upload_date, .id, .duration, .title] | @tsv' | sort 335 336 # only uploads inside a window, which is how you tie output to an event 337 yt-dlp --flat-playlist --dateafter 20260101 --datebefore 20260331 \ 338 -O '%(upload_date)s %(id)s %(title)s' 'https://youtube.com/@channel/videos' 339 340 # subtitles for every video in a playlist, as text, no media 341 yt-dlp --write-auto-subs --sub-langs 'en.*' --convert-subs srt --skip-download \ 342 -o '%(upload_date)s-%(id)s.%(ext)s' 'https://youtube.com/playlist?list=PLID' 343 344 # then the actual search, across the lot 345 grep -rin 'phrase you care about' *.srt 346 ``` 347 348 `--flat-playlist` is what keeps a channel listing cheap: it reads the playlist entries without 349 resolving each video, so a 2,000-video channel is one request chain rather than 2,000. The price is 350 that some fields are absent from flat entries, so confirm a field exists in `--dump-json` output 351 before you build a pipeline on it. 352 353 ## Reddit 354 355 - Profiles, posts and comment trees are readable without an account, and appending `.json` to any 356 Reddit URL returns the same thing as structured data. 357 - Deleted and removed content is the point of interest, and Reddit does not serve it. The 358 historical archives do — Arctic Shift below is the live successor to Pushshift. 359 - Subreddit moderator lists, wiki pages and automoderator configs are public and frequently 360 forgotten by the people who wrote them. 361 362 The `.json` trick, with the parameters that matter: 363 364 ```bash 365 # any Reddit URL plus .json, with escaping turned off and a full page 366 curl -s -A 'research-tooling/1.0 (contact: you@example.test)' \ 367 'https://www.reddit.com/user/some_user/comments/.json?limit=100&raw_json=1' \ 368 | jq -r '.data.children[] | "\(.data.created_utc) \(.data.subreddit) \(.data.body[0:100])"' 369 370 # page with the fullname the previous response handed back 371 curl -s -A 'research-tooling/1.0' \ 372 'https://www.reddit.com/r/subreddit/new/.json?limit=100&raw_json=1&after=t3_abc123' 373 ``` 374 375 `raw_json=1` is not cosmetic: without it Reddit replaces `<`, `>` and `&` with HTML entities in 376 every string field, so a URL or a quoted snippet comes back mangled and your `grep` misses it. 377 Expect this route to fail anyway. Reddit now returns **403 to unauthenticated JSON requests from 378 datacentre address ranges regardless of User-Agent** — tested from a hosted address, a custom agent 379 string makes no difference — so a VPS or CI runner gets nothing while a residential browser 380 session gets everything. That asymmetry is exactly why the archive below is not optional. 381 382 ## Bluesky 383 384 Structurally the most open platform here, because the AT Protocol is a public API by design: every 385 post, follow and profile is a record in a repository you can read without an account. 386 387 - **DIDs are permanent.** A handle is a pointer; the `did:plc:…` identifier survives every rename, 388 and the PLC directory keeps an audit log of every change to it — including the account's first 389 operation, which dates its creation. 390 - **Follow graphs are fully enumerable**, which makes network analysis possible in a way it is not 391 anywhere else on this page. 392 - **Search is the one thing that is not open.** Post search now needs an authenticated session; 393 profile and feed reads do not. 394 395 There are two layers and it is worth knowing which one you are on. The AppView 396 (`public.api.bsky.app`) serves the rendered, moderated view that the app shows. The account's own 397 PDS serves the raw repository, which includes record types the app never displays: 398 399 ```bash 400 # which record collections this account actually holds -- not just posts 401 curl -s 'https://bsky.social/xrpc/com.atproto.repo.describeRepo?repo=atproto.com' \ 402 | jq '{handle, did, handleIsCorrect, collections}' 403 404 # raw records from one collection, straight out of the repository 405 curl -s 'https://bsky.social/xrpc/com.atproto.repo.listRecords?repo=atproto.com&collection=app.bsky.feed.post&limit=100&reverse=true' \ 406 | jq -r '.records[] | "\(.value.createdAt) \(.uri | split("/") | last) \(.value.text[0:80])"' 407 ``` 408 409 `reverse=true` on `listRecords` gives you the **oldest** records first, which is how you read an 410 account's first ever posts without paging through everything since. The full tool coverage, the 411 AppView endpoints and the `goat` CLI are below under [Key tools](#key-tools). 412 413 ## Key tools 414 415 ### yt-dlp 416 417 Does far more than download: it reads the platform's own metadata for a video, a channel or a 418 playlist, and it covers well over a thousand sites, so the same invocation works on YouTube, 419 TikTok, Twitch, Instagram and most of the long tail. For dating footage and enumerating a 420 channel's output it is the first tool to reach for. 421 422 ```bash 423 pipx install yt-dlp 424 425 # everything the platform will tell you about one video, without downloading it 426 yt-dlp --dump-json 'https://youtube.com/watch?v=VIDEOID' | jq '{upload_date,uploader_id,duration,view_count}' 427 428 # just the fields you need, printed — faster to read than the full JSON 429 yt-dlp -O '%(upload_date)s %(uploader_id)s %(title)s' 'https://youtube.com/watch?v=VIDEOID' 430 431 # the whole channel as a listing, without touching the videos themselves 432 yt-dlp --flat-playlist --dump-json 'https://youtube.com/@channel/videos' \ 433 | jq -r '[.upload_date,.id,.title] | @tsv' 434 435 # the entire playlist as ONE json object, for feeding a script rather than a reader 436 yt-dlp --flat-playlist --dump-single-json 'https://youtube.com/@channel/videos' \ 437 | jq '{channel: .uploader, count: (.entries | length)}' 438 439 # auto-generated subtitles, which turn hours of video into greppable text 440 yt-dlp --write-auto-subs --sub-langs en --skip-download 'https://youtube.com/watch?v=VIDEOID' 441 442 # which subtitle tracks exist at all, before you guess a language code 443 yt-dlp --list-subs 'https://youtube.com/watch?v=VIDEOID' 444 445 # comments, saved alongside the metadata, for identifying participants 446 yt-dlp --write-info-json --write-comments --skip-download 'https://youtube.com/watch?v=VIDEOID' 447 448 # the thumbnail set, which is a reverse-image-search input and sometimes an earlier frame 449 yt-dlp --write-thumbnail --skip-download 'https://youtube.com/watch?v=VIDEOID' 450 451 # one clip out of a long stream, by timestamp, instead of the whole six hours 452 yt-dlp --download-sections '*01:12:30-01:14:00' 'https://youtube.com/watch?v=VIDEOID' 453 454 # a date window across a channel, for tying output to an event 455 yt-dlp --flat-playlist --dateafter 20260101 --datebefore 20260331 \ 456 -O '%(upload_date)s %(id)s' 'https://youtube.com/@channel/videos' 457 458 # filter on any metadata field rather than on dates; "&" is how conditions AND together 459 yt-dlp --match-filters 'duration > 600 & view_count > 10000' \ 460 -O '%(id)s %(duration)s %(view_count)s' 'https://youtube.com/@channel/videos' 461 462 # "?" after the operator keeps items whose field is absent, instead of dropping them silently 463 yt-dlp --match-filters 'view_count >? 10000' \ 464 -O '%(id)s %(view_count)s' 'https://youtube.com/@channel/videos' 465 466 # a TikTok video with its metadata — same tool, same flags 467 yt-dlp --write-info-json 'https://www.tiktok.com/@user/video/1234567890123456789' 468 469 # items 1 to 50 of a playlist, with an archive file so a resumed run does not re-fetch 470 yt-dlp -I 1:50 --download-archive seen.txt 'https://youtube.com/playlist?list=PLID' 471 472 # a livestream from its start, for an event already in progress 473 yt-dlp --live-from-start 'https://youtube.com/watch?v=LIVEID' 474 475 # the logged-in surface, when a public one does not exist — uses your browser's session 476 yt-dlp --cookies-from-browser firefox --dump-json URL 477 ``` 478 479 `upload_date` is the platform's own value and is a reliable earliest-possible date for footage; 480 it is not the date the footage was shot. Fields vary by extractor, so check with `--dump-json` 481 before scripting against a key — `--match-filters` silently matches nothing when the field you 482 named does not exist on that extractor, which looks identical to a channel with no matching 483 videos. The `>?` form above is the documented escape: it treats an absent field as a pass, so a 484 sweep that returns nothing is telling you about the videos rather than about the extractor. Passing `--cookies-from-browser` attaches your real session to every request, which 485 attributes the activity to that account — use a research profile. Extractors break when platforms 486 change, so `pipx upgrade yt-dlp` before blaming the URL. 487 488 ### gallery-dl 489 490 The image-and-gallery counterpart to yt-dlp: it walks a profile, a tag or a thread and takes the 491 media plus the per-item JSON. Where it earns its place is bulk profile capture with metadata 492 intact — the captions, timestamps and IDs that make a media set evidence rather than a folder of 493 pictures. 494 495 ```bash 496 pipx install gallery-dl 497 498 # what it would fetch, and the metadata for each item, without downloading 499 gallery-dl -j 'https://www.instagram.com/username/' 500 501 # the available metadata keys for a URL, so your filters reference real fields 502 gallery-dl -K 'https://www.reddit.com/user/username/submitted/' 503 504 # media plus a JSON sidecar per file, into a named directory 505 gallery-dl --write-metadata -D ./capture/username 'https://twitter.com/username/media' 506 507 # direct URLs only, to hand to an archiving tool instead of downloading yourself 508 gallery-dl -g 'https://www.reddit.com/r/subreddit/comments/abc123/' 509 510 # a chosen field per item, printed as the run goes, instead of parsing json afterwards 511 gallery-dl --no-download -N '{date} {num} {filename}' 'https://www.reddit.com/user/username/submitted/' 512 513 # the first 50 items of a long profile, which is usually all you need to characterise it 514 gallery-dl --range 1-50 'https://www.flickr.com/photos/username/' 515 516 # only the large images, skipping avatars and thumbnails 517 gallery-dl --filter 'image_width >= 1000' 'https://example-gallery.test/user/x' 518 519 # an archive file, so re-running a monitored profile fetches only what is new 520 gallery-dl --download-archive seen.sqlite3 'https://www.instagram.com/username/' 521 522 # many targets from a file, which is how a watchlist gets run 523 gallery-dl -i targets.txt --download-archive seen.sqlite3 524 525 # slow and polite: the alternative is a 429 and then a block 526 gallery-dl --sleep 3 --sleep-request 2 --sleep-429 120 'https://www.instagram.com/username/' 527 528 # hash each file as it lands, so the capture is checkable later 529 gallery-dl --exec 'sha256sum {} >> hashes.txt' -D ./capture 'https://twitter.com/username/media' 530 531 # keep the raw HTML the extractor parsed, which is the only way to debug an empty run 532 gallery-dl --write-pages -j 'https://www.instagram.com/username/' 533 534 # a logged-in session where the public surface is gone 535 gallery-dl --cookies-from-browser firefox 'https://www.instagram.com/username/' 536 ``` 537 538 Per-site options live in a config file (`gallery-dl --config-create` writes a starter), and most 539 rate-limit problems are solved there rather than on the command line — `-o KEY=VALUE` sets one 540 inline when you do not want to edit the file. Timestamps in the sidecar JSON are the platform's, 541 which is the point — but the field name differs per extractor, so read `-K` output before you build 542 a timeline. A profile capture is a snapshot: items deleted before you ran it are simply absent, and 543 nothing in the output tells you that. An extractor that returns zero items is far more often broken 544 than correct, and `--write-pages` is how you tell the difference. 545 546 ### Instaloader 547 548 Instagram specifically, and still working where most Instagram tooling has died. Public profiles, 549 hashtags, locations and stories, with geotags and comments where they exist. 550 551 ```bash 552 pipx install instaloader 553 554 # a public profile: posts, captions and metadata JSON 555 instaloader profile_name 556 557 # metadata only — no media, much faster, much less conspicuous 558 instaloader --no-pictures --no-videos profile_name 559 560 # comments and geotags, which are the parts worth having 561 instaloader --comments --geotags profile_name 562 563 # just the profile picture at full resolution, for reverse image search 564 instaloader --no-posts profile_name 565 566 # cap a location run; --count does not apply to ordinary profile targets 567 instaloader --count 50 --no-videos '%212998617' 568 569 # posts within a date window, using a Python filter expression 570 instaloader --post-filter 'date_utc >= datetime(2026,1,1)' profile_name 571 572 # only posts that carry a location, which is the geolocation-relevant subset 573 instaloader --post-filter 'location is not None' --geotags profile_name 574 575 # a hashtag rather than an account 576 instaloader '#hashtag' 577 578 # a numeric location id, which is what instagram-location-search below produces 579 instaloader '%212998617' 580 581 # everywhere this account has been tagged by other people 582 instaloader --tagged --no-posts profile_name 583 584 # one post by shortcode; the leading dash needs the -- separator first 585 instaloader -- -B_Nmp6MlQGV 586 587 # readable JSON instead of xz-compressed, for grepping a capture afterwards 588 instaloader --no-compress-json --no-pictures --no-videos profile_name 589 590 # stories and highlights need a session; both attribute the view to that account 591 instaloader --login YOUR_RESEARCH_ACCOUNT --stories --highlights profile_name 592 593 # reuse a browser session instead of handing credentials to the tool 594 instaloader --load-cookies firefox --sessionfile ./research.session profile_name 595 596 # stop on the first rate-limit response rather than grinding into a block 597 instaloader --abort-on 429,401 --max-connection-attempts 1 profile_name 598 599 # resume by stopping at the first already-downloaded post 600 instaloader --fast-update profile_name 601 602 # resume by recorded timestamp instead, which is the only resume a metadata-only run can use 603 instaloader --no-pictures --no-videos --latest-stamps stamps.ini profile_name 604 ``` 605 606 **Viewing a story is attributed.** The account whose session you used appears in the poster's 607 viewer list, so `--stories` is never a passive operation. Unauthenticated use is rate-limited 608 aggressively and then blocked by IP; authenticated use gets the account flagged and sometimes 609 disabled — `--abort-on 429` is the difference between one bad response and a dead account. 610 `--no-pictures` cannot be combined with `--fast-update`, because the resume logic keys off 611 downloaded media, so a metadata-only monitoring run needs `--latest-stamps` instead. Geotags are 612 only present where the poster added them, and the location name is user-chosen, not GPS. 613 614 ### instagram-location-search 615 616 Bellingcat's answer to "what is Instagram's own name for this place". It takes a coordinate and 617 returns the location tags Instagram knows about nearby, with their numeric IDs — which is the 618 pivot, because an ID is an addressable target for Instaloader while a place name is not. 619 620 ```bash 621 pip install instagram-location-search 622 623 # a coordinate to a list of location tags, as CSV 624 instagram-location-search --cookie "$IG_COOKIE" --lat 51.5074 --lng -0.1278 --csv locs.csv 625 626 # every output format at once: the raw API response, a geospatial layer, and a map to eyeball 627 instagram-location-search --cookie "$IG_COOKIE" --lat 51.5074 --lng -0.1278 \ 628 --json locs.json --geojson locs.geojson --map locs.html 629 630 # just the IDs, which is the form the next tool wants 631 instagram-location-search --cookie "$IG_COOKIE" --lat 51.5074 --lng -0.1278 --ids locs.txt 632 633 # then feed each one to Instaloader as a location target 634 while read -r id; do instaloader --no-videos --geotags "%$id"; done < locs.txt 635 ``` 636 637 It needs a session cookie, so every search is attributed to that account. The locations it returns 638 are *Instagram's* places, which are user-created: there are duplicates, misspellings, joke entries 639 and venues that have closed, and a radius search returns them by proximity to the coordinate rather 640 than by relevance. A location tag on a post means the poster selected that place from a list, not 641 that a GPS fix put them there — see [Geolocation & Chronolocation](/sheets/osint/geolocation) for 642 what actually constitutes placing an image. 643 644 ### Telegram: Telepathy and Telethon 645 646 [Telepathy](https://github.com/proseltd/Telepathy-Community) is the toolkit the Telegram sections 647 of older OSINT guides point at, but it has been **unmaintained since July 2024** and the current 648 source branch does not install cleanly. The commands below document the published 2.3.4 interface; 649 for anything you need to keep working, the durable route is Telethon, which tracks the Telegram 650 API itself. 651 652 ```bash 653 pip3 install telepathy 654 655 telepathy -t channelname # basic channel scan 656 telepathy -t channelname -c # comprehensive: adds the message history archive 657 telepathy -t channelname -c -f # plus a forward edgelist, the useful part 658 telepathy -t channelname -c -m # media archiving alongside the comprehensive scan 659 telepathy -t channelname -c -r # archive channel replies as well as posts 660 telepathy -t username -u # look up a user your account has encountered 661 telepathy -e # export your account's chat list to CSV 662 telepathy -t channelname -a 2 # run from an alternative number or API key set 663 ``` 664 665 Both need API credentials from [my.telegram.org](https://my.telegram.org/), which are tied to a 666 real phone number — use one you are willing to lose. Telethon in forty lines does the part that 667 matters, and does not rot: 668 669 ```python 670 # pip install telethon 671 from collections import Counter 672 from telethon.sync import TelegramClient 673 from telethon.errors import FloodWaitError 674 from telethon.tl.functions.channels import GetFullChannelRequest 675 from telethon.tl.types import InputMessagesFilterPhotos 676 677 with TelegramClient('research-session', API_ID, API_HASH) as client: 678 entity = client.get_entity('channelname') 679 print(entity.id, entity.title, entity.date) # channel id and creation date 680 681 # what Telegram itself reports about the channel, including the linked discussion group 682 full = client(GetFullChannelRequest(entity)).full_chat 683 print(full.participants_count, full.linked_chat_id, (full.about or '')[:120]) 684 685 # message history, newest first; message IDs are sequential, so gaps are deletions 686 for msg in client.iter_messages(entity, limit=500): 687 fwd = msg.forward.chat.username if msg.forward and msg.forward.chat else None 688 print(msg.id, msg.date.isoformat(), fwd or '-', (msg.message or '')[:80]) 689 690 # the count Telegram reports, against the newest id: the difference is deleted content 691 page = client.get_messages(entity, limit=100) 692 newest = page[0] 693 print('reported total', page.total, 'newest id', newest.id, 'gap', newest.id - page.total) 694 695 # oldest first, which is how you read a channel's first posts without paging the lot 696 for msg in client.iter_messages(entity, limit=20, reverse=True): 697 print(msg.id, msg.date.isoformat(), (msg.message or '')[:80]) 698 699 # server-side keyword search inside one channel -- far cheaper than fetching everything 700 for msg in client.iter_messages(entity, search='keyword', limit=100): 701 print(msg.id, msg.date.isoformat(), (msg.message or '')[:120]) 702 703 # only messages carrying photos, for building a verification set 704 for msg in client.iter_messages(entity, filter=InputMessagesFilterPhotos, limit=50): 705 print(msg.id, msg.date.isoformat(), msg.file.name if msg.file else '-') 706 707 # everything one account posted in a group, which is the behavioural view 708 for msg in client.iter_messages('groupname', from_user='username', limit=200): 709 print(msg.id, msg.date.isoformat(), (msg.message or '')[:80]) 710 711 # the forward graph: which channels this one amplifies, with counts 712 sources = Counter( 713 m.forward.chat.username 714 for m in client.iter_messages(entity, limit=2000) 715 if m.forward and m.forward.chat 716 ) 717 print(sources.most_common(20)) 718 719 # membership, where the group exposes it — you appear in their list too 720 for user in client.iter_participants(entity, limit=200): 721 print(user.id, user.username, user.first_name) 722 ``` 723 724 Sequential message IDs are the quiet win here: a gap between IDs is deleted content, and the first 725 ID dates the channel. The `search=` parameter runs server-side, which matters for a channel with 726 100,000 messages — it is the difference between one request and a thousand. Joining a group to read 727 it adds your research account to a list other members can see, and scraping at speed from one 728 account is exactly the behaviour Telegram bans for. Rate limits surface as `FloodWaitError` with a 729 `seconds` attribute — sleep for exactly that, rather than retrying, because retrying extends it. 730 731 ### telegram-phone-number-checker 732 733 Bellingcat's tool for the one Telegram question with a clean answer: is this phone number 734 registered, and if so what does the account show. Also takes usernames. 735 736 ```bash 737 pip install telegram-phone-number-checker 738 739 # credentials go in a .env: API_ID, API_HASH, PHONE_NUMBER 740 telegram-phone-number-checker --phone-numbers +14155550123 741 742 # several at once, comma-separated 743 telegram-phone-number-checker --phone-numbers +14155550123,+14155550124 744 745 # pull the profile photos too, for reverse image search 746 telegram-phone-number-checker --phone-numbers +14155550123 --download-profile-photos 747 748 # by username instead of number 749 telegram-phone-number-checker --usernames johndoe 750 751 # both in one run, to a named output file 752 telegram-phone-number-checker --phone-numbers +14155550123 --usernames johndoe --output results.json 753 754 # credentials on the command line, overriding the .env, for a one-off run on a second account 755 telegram-phone-number-checker --api-id 123456 --api-hash abcdef0123456789 \ 756 --api-phone-number +14155550000 --phone-numbers +14155550123 757 758 # through a proxy, since the lookups come from your account 759 pip install 'telegram-phone-number-checker[proxy]' 760 telegram-phone-number-checker --phone-numbers +14155550123 --proxy socks5://127.0.0.1:1080 761 ``` 762 763 Checking a number works by adding it to your account's contacts, which is a write operation on 764 Telegram's side: the README's own advice is not to use a personal account, and that is not 765 boilerplate. A negative result means the number is not registered *or* the account's privacy 766 settings hide it from non-contacts — the tool cannot tell you which. See 767 [email and phone OSINT](/sheets/osint/email-and-phone) for the rest of the number work. 768 769 ### tiktok-hashtag-analysis 770 771 Bellingcat's bulk primitive for TikTok: collect the posts carrying one or more hashtags, store the 772 metadata, and chart which hashtags co-occur. That co-occurrence graph is how a coordinated 773 campaign becomes visible, which individual video analysis never shows. 774 775 ```bash 776 pip install tiktok-hashtag-analysis 777 778 # collect posts for one or more hashtags 779 tiktok-hashtag-analysis london paris 780 781 # hashtags from a file, when the list is long 782 tiktok-hashtag-analysis --file hashtags.txt --output-dir ./tiktok-run 783 784 # the top 20 co-occurring hashtags, as a table 785 tiktok-hashtag-analysis london --number 20 --table 786 787 # the same, plotted 788 tiktok-hashtag-analysis london --number 20 --plot 789 790 # download the videos too, capped per hashtag 791 tiktok-hashtag-analysis london --download --limit 200 792 793 # short forms, for the same thing: -d download, -t table, -p plot 794 tiktok-hashtag-analysis london -d -t --number 20 --limit 50 795 796 # keep a log file, which is the only record of what a long run actually did 797 tiktok-hashtag-analysis london --log ./run.log --output-dir ./tiktok-run 798 799 # watch the browser work, which is how you debug a run that returns nothing 800 tiktok-hashtag-analysis london --headed -v 801 ``` 802 803 It drives a real browser, so it is slow and it breaks when TikTok changes its front end — a run 804 that returns zero posts for a busy hashtag means the scraper is broken, not that the hashtag is 805 empty. Check with `--headed` before concluding anything. Hashtag collection is a sample, never the 806 complete set, so counts are comparative at best, and two runs an hour apart will not agree. For a 807 single video's exact upload time, decode the ID as shown under [TikTok](#tiktok) above, and 808 `yt-dlp --dump-json` gets the rest. 809 810 ### Arctic Shift 811 812 The live successor to Pushshift for Reddit: a public, keyless archive of posts and comments, 813 including content since deleted from Reddit itself. This is the only practical way to read a 814 removed comment or reconstruct a user's history after they wiped it. 815 816 ```bash 817 AS='https://arctic-shift.photon-reddit.com/api' 818 819 # everything an author posted, oldest first 820 curl -s -G "$AS/posts/search" --data-urlencode 'author=some_user' \ 821 --data-urlencode 'sort=asc' --data-urlencode 'limit=100' \ 822 | jq -r '.data[] | "\(.created_utc) \(.subreddit) \(.title)"' 823 824 # their comments, which are usually more revealing than their posts 825 curl -s -G "$AS/comments/search" --data-urlencode 'author=some_user' \ 826 --data-urlencode 'limit=100' | jq -r '.data[] | "\(.subreddit) \(.body[0:120])"' 827 828 # keyword search inside a subreddit and a date window 829 curl -s -G "$AS/posts/search" --data-urlencode 'subreddit=worldnews' \ 830 --data-urlencode 'title=wuhan' --data-urlencode 'after=2019-12-30' --data-urlencode 'limit=10' 831 832 # title and selftext together, which is what 'query' is for 833 curl -s -G "$AS/posts/search" --data-urlencode 'subreddit=osint' \ 834 --data-urlencode 'query=telegram' --data-urlencode 'limit=25' 835 836 # only the fields you want, which keeps a wide sweep manageable 837 curl -s -G "$AS/posts/search" --data-urlencode 'subreddit=osint' \ 838 --data-urlencode 'fields=title,created_utc,author' --data-urlencode 'limit=100' 839 840 # posts linking a specific domain — how a site spreads across subreddits 841 curl -s -G "$AS/posts/search" --data-urlencode 'url=example-news-daily.com' \ 842 --data-urlencode 'limit=100' | jq -r '.data[].subreddit' | sort | uniq -c | sort -rn 843 844 # posting cadence rather than content: a month-by-month count 845 curl -s -G "$AS/posts/search/aggregate" --data-urlencode 'subreddit=osint' \ 846 --data-urlencode 'aggregate=created_utc' --data-urlencode 'frequency=month' \ 847 --data-urlencode 'after=2026-01-01' | jq -r '.data[] | "\(.created_utc[0:7]) \(.count)"' 848 849 # where an account actually operates, weighted and ranked 850 curl -s -G "$AS/users/interactions/subreddits" --data-urlencode 'author=some_user' \ 851 --data-urlencode 'limit=10' | jq -r '.data[] | "\(.count)\t\(.subreddit)"' 852 853 # who an account talks to, which is the social graph Reddit does not expose 854 curl -s -G "$AS/users/interactions/users" --data-urlencode 'author=some_user' \ 855 --data-urlencode 'min_count=3' --data-urlencode 'limit=20' 856 857 # every wiki page a subreddit has, including the ones nobody links to 858 curl -s -G "$AS/subreddits/wikis/list" --data-urlencode 'subreddit=osint' | jq -r '.data[]' 859 860 # the full comment tree under one post, up to 25,000 comments 861 curl -s -G "$AS/comments/tree" --data-urlencode 'link_id=abc123' \ 862 --data-urlencode 'limit=9999' 863 864 # known IDs straight to records, up to 500 per call 865 curl -s -G "$AS/comments/ids" --data-urlencode 'ids=t1_aaa,t1_bbb' 866 867 # accounts by username prefix, for a sockpuppet family 868 curl -s -G "$AS/users/search" --data-urlencode 'author_prefix=civicwatch' \ 869 --data-urlencode 'sort_type=author' --data-urlencode 'limit=100' | jq -r '.data[].author' 870 871 # an RSS feed of a standing query, for monitoring rather than a one-off sweep 872 curl -s -G "$AS/comments/search" --data-urlencode 'subreddit=osint' \ 873 --data-urlencode 'format=rss' --data-urlencode 'limit=25' 874 ``` 875 876 A real response, from the aggregate query above: 877 878 ```json 879 {"data":[{"created_utc":"2025-12-31T23:00:00.000Z","count":"390"}, 880 {"created_utc":"2026-01-31T23:00:00.000Z","count":"376"}, 881 {"created_utc":"2026-02-28T23:00:00.000Z","count":"503"}, 882 {"created_utc":"2026-03-31T22:00:00.000Z","count":"450"}]} 883 ``` 884 885 Note the bucket boundaries: `23:00:00Z` then `22:00:00Z` after March, because the buckets are cut 886 in a local timezone that observes DST, so a month's `created_utc` label is the start of the bucket 887 and not the month. Keyword search is constrained by design — `title` works only alongside `author` 888 or `subreddit`, `body` only alongside `author`, `subreddit`, `link_id` or `parent_id` — so a bare 889 keyword sweep is refused, and it is refused for very active authors and subreddits too. The archive 890 holds what it ingested at the time, so an edit after ingestion is invisible and a post removed 891 within seconds of being made may never have been captured. A record here is evidence that text was 892 published, not that it is still live, and the live check is a separate step. 893 894 ### Bluesky and the AT Protocol 895 896 No key, no account, and a genuinely public API — the most researcher-friendly surface of any 897 current platform. Everything below runs against the public AppView except post search, which now 898 requires an authenticated session. 899 900 ```bash 901 API='https://public.api.bsky.app/xrpc' 902 903 # handle to DID: the identifier that survives renames 904 curl -s "$API/com.atproto.identity.resolveHandle?handle=bellingcat.com" | jq -r .did 905 906 # the PLC audit log: every handle and server change, and the creation date in entry one 907 curl -s 'https://plc.directory/did:plc:z72i7hdynmk6r22z27h6tvur/log/audit' \ 908 | jq -r '.[] | "\(.createdAt) \(.operation.alsoKnownAs // [] | join(","))"' 909 910 # profile: display name, counts, and the account's own description 911 curl -s "$API/app.bsky.actor.getProfile?actor=bellingcat.com" \ 912 | jq '{handle,displayName,followersCount,postsCount,createdAt}' 913 914 # up to 25 accounts in one request, which is how you compare a suspected cluster 915 curl -s -G "$API/app.bsky.actor.getProfiles" \ 916 --data-urlencode 'actors=bellingcat.com' --data-urlencode 'actors=atproto.com' \ 917 | jq -r '.profiles[] | "\(.createdAt) \(.followersCount)\t\(.handle)"' 918 919 # the author's posts, paged with the cursor the response returns 920 curl -s "$API/app.bsky.feed.getAuthorFeed?actor=bellingcat.com&limit=100" \ 921 | jq -r '.feed[] | "\(.post.indexedAt) \(.post.record.text[0:100])"' 922 923 # only posts carrying media, which is the usual filter for verification work 924 curl -s "$API/app.bsky.feed.getAuthorFeed?actor=bellingcat.com&limit=50&filter=posts_with_media" 925 926 # the account's own threads without the replies noise 927 curl -s "$API/app.bsky.feed.getAuthorFeed?actor=bellingcat.com&limit=50&filter=posts_no_replies" 928 929 # the follow graph, fully enumerable in both directions 930 curl -s "$API/app.bsky.graph.getFollows?actor=bellingcat.com&limit=100" | jq -r '.follows[].handle' 931 curl -s "$API/app.bsky.graph.getFollowers?actor=bellingcat.com&limit=100" | jq -r '.followers[].handle' 932 933 # the lists an account curates, which say who it considers relevant 934 curl -s "$API/app.bsky.graph.getLists?actor=bellingcat.com&limit=50" \ 935 | jq -r '.lists[] | "\(.purpose)\t\(.listItemCount)\t\(.name)"' 936 937 # who amplified one post, and who liked it -- both need the post's AT-URI 938 curl -s "$API/app.bsky.feed.getRepostedBy?uri=at://did:plc:xxxx/app.bsky.feed.post/yyyy&limit=100" \ 939 | jq -r '.repostedBy[].handle' 940 curl -s "$API/app.bsky.feed.getLikes?uri=at://did:plc:xxxx/app.bsky.feed.post/yyyy&limit=100" \ 941 | jq -r '.likes[] | "\(.createdAt) \(.actor.handle)"' 942 943 # a whole thread, replies included, from one post's AT URI 944 curl -s "$API/app.bsky.feed.getPostThread?uri=at://did:plc:xxxx/app.bsky.feed.post/yyyy" 945 946 # post search needs a session — this returns an auth error on the public host 947 curl -s -G "$API/app.bsky.feed.searchPosts" --data-urlencode 'q=osint' \ 948 --data-urlencode 'sort=latest' --data-urlencode 'since=2026-01-01' --data-urlencode 'limit=25' 949 ``` 950 951 Once you have a session, `searchPosts` is the richest query surface on the platform: `q` plus 952 `author`, `mentions`, `domain`, `url`, `lang`, repeated `tag` parameters that AND together, and 953 `since`/`until` which take a bare `YYYY-MM-DD`. Authenticate with 954 `com.atproto.server.createSession` against the account's own PDS and send the returned 955 `accessJwt` as a bearer token: 956 957 ```bash 958 # a session on the account's own PDS; use an app password, never the real one 959 PDS='https://bsky.social' 960 TOKEN=$(curl -s -X POST "$PDS/xrpc/com.atproto.server.createSession" \ 961 -H 'Content-Type: application/json' \ 962 -d '{"identifier":"you.bsky.social","password":"xxxx-xxxx-xxxx-xxxx"}' | jq -r .accessJwt) 963 964 # every post from one account mentioning a domain, inside a date window 965 curl -s -G "$PDS/xrpc/app.bsky.feed.searchPosts" -H "Authorization: Bearer $TOKEN" \ 966 --data-urlencode 'q=*' --data-urlencode 'author=bellingcat.com' \ 967 --data-urlencode 'domain=example-news-daily.com' --data-urlencode 'since=2026-01-01' \ 968 | jq -r '.posts[] | "\(.record.createdAt) \(.author.handle) \(.record.text[0:100])"' 969 970 # posts carrying two tags at once, which is the coordination signal 971 curl -s -G "$PDS/xrpc/app.bsky.feed.searchPosts" -H "Authorization: Bearer $TOKEN" \ 972 --data-urlencode 'q=*' --data-urlencode 'tag=osint' --data-urlencode 'tag=verification' \ 973 --data-urlencode 'sort=latest' | jq -r '.posts[].author.handle' | sort | uniq -c | sort -rn 974 ``` 975 976 The PLC audit log is the find here: it is an append-only record of every handle the account has 977 used and every server it has moved between, dated, with the first entry establishing when the 978 account was created. `indexedAt` is when the relay saw the post, not when the author wrote it — 979 `record.createdAt` is the author's claim and is trivially spoofable, as is the TID in the record 980 key, so quote both and say which is which. `searchPosts` is explicitly documented as possibly 981 non-public, so treat the public host returning results for it as luck rather than a contract, and 982 note that `since`/`until` filter on the server's `sortAt`, which may not match `createdAt` at all. 983 Authenticating turns a keyless read into an attributable one. 984 985 ### goat 986 987 The reference AT Protocol CLI, from Bluesky themselves. It does the things `curl` makes awkward: 988 exporting a whole repository as a single file, decoding record keys, walking the PLC operation log, 989 and tailing the firehose. For Bluesky work it replaces about ten of the commands above. 990 991 ```bash 992 brew install goat # or: go install github.com/bluesky-social/goat@latest 993 994 # resolve an identity to its DID document, including which PDS holds the data 995 goat resolve wyden.senate.gov 996 997 # every record collection the account holds -- including ones the app never renders 998 goat ls -c dril.bsky.social 999 1000 # one record, as the author wrote it 1001 goat get at://dril.bsky.social/app.bsky.feed.post/3kkreaz3amd27 1002 1003 # the entire repository as a CAR file: every record, in one evidential artefact 1004 goat repo export jay.bsky.team 1005 1006 # and the blobs, which is every image the account has ever uploaded 1007 goat blob export jay.bsky.team 1008 1009 # the full identity history: handle changes, PDS moves, key rotations, dated 1010 goat plc history atproto.com 1011 1012 # a record key back to the time the client claims it was written 1013 goat syntax tid inspect 3kzifvcppte22 1014 1015 # live posts matching nothing in particular, for watching an event unfold 1016 goat firehose --ops -c app.bsky.feed.post 1017 1018 # handle changes across the whole network, which is a renaming-detection feed 1019 goat firehose --account-events | jq .payload.handle 1020 1021 # any XRPC query against any host, which is the escape hatch 1022 goat xrpc query https://public.api.bsky.app app.bsky.actor.getProfile actor==atproto.com 1023 ``` 1024 1025 `goat repo export` is the one to reach for when the account might be deleted: a CAR file is the 1026 whole repository at one moment, signed, and it keeps working after the account does not. Note the 1027 `==` in the `xrpc query` form — that is how `goat` distinguishes a query parameter from a header, 1028 and a single `=` there means something else. Reads need no authentication; `goat account login` 1029 exists for writes and for the handful of authenticated reads, and logging in attributes everything 1030 after it. A TID decoded by `syntax tid inspect` reads the clock of whatever generated the record 1031 key, not the network's, and the TID specification notes that known TIDs can be reused — so it 1032 dates a claim rather than an event. 1033 1034 ### Meta Ad Library API 1035 1036 The one Meta surface that is genuinely open, and the only reliable way to get spend and reach data 1037 for political and issue advertising. Searchable by keyword or page, with the creative, the dates, 1038 the platforms and the demographic breakdown. 1039 1040 ```bash 1041 # access needs a Meta developer app, ID verification and Ad Library API approval first 1042 TOKEN="$META_TOKEN" 1043 G='https://graph.facebook.com/v23.0/ads_archive' 1044 1045 # keyword search; ad_reached_countries is mandatory, and so is search_terms or search_page_ids 1046 curl -s -G "$G" -d "access_token=$TOKEN" \ 1047 -d 'search_terms=climate' -d "ad_reached_countries=['GB']" \ 1048 -d 'fields=page_name,ad_delivery_start_time,ad_snapshot_url' | jq -r '.data[].page_name' | sort -u 1049 1050 # an exact phrase rather than unordered keywords 1051 curl -s -G "$G" -d "access_token=$TOKEN" \ 1052 -d 'search_terms=net zero' -d 'search_type=KEYWORD_EXACT_PHRASE' \ 1053 -d "ad_reached_countries=['GB']" -d 'fields=page_name,ad_creative_bodies' 1054 1055 # one page's entire ad history, which is the attribution-relevant query 1056 curl -s -G "$G" -d "access_token=$TOKEN" \ 1057 -d 'search_page_ids=123456789' -d "ad_reached_countries=['GB']" \ 1058 -d 'ad_active_status=ALL' \ 1059 -d 'fields=ad_creative_bodies,ad_delivery_start_time,ad_delivery_stop_time,spend,impressions' 1060 1061 # who paid, which is the field the transparency rules exist for 1062 curl -s -G "$G" -d "access_token=$TOKEN" \ 1063 -d 'search_terms=election' -d "ad_reached_countries=['GB']" \ 1064 -d 'fields=bylines,beneficiary_payers,page_name,spend,currency,publisher_platforms' 1065 1066 # where it ran and to whom 1067 curl -s -G "$G" -d "access_token=$TOKEN" \ 1068 -d 'search_terms=election' -d "ad_reached_countries=['GB']" \ 1069 -d 'fields=delivery_by_region,demographic_distribution,languages' 1070 1071 # only the ads that ran as video, on Instagram rather than Facebook 1072 curl -s -G "$G" -d "access_token=$TOKEN" \ 1073 -d 'search_terms=election' -d "ad_reached_countries=['GB']" \ 1074 -d 'media_type=VIDEO' -d "publisher_platforms=['INSTAGRAM']" \ 1075 -d 'fields=page_name,ad_snapshot_url,publisher_platforms' 1076 1077 # the non-political categories, which most people never query 1078 curl -s -G "$G" -d "access_token=$TOKEN" \ 1079 -d 'search_terms=apartment' -d "ad_reached_countries=['GB']" \ 1080 -d 'ad_type=HOUSING_ADS' -d 'fields=page_name,ad_creative_bodies,ad_delivery_start_time' 1081 1082 # a date window, for tying a campaign to an event 1083 curl -s -G "$G" -d "access_token=$TOKEN" \ 1084 -d 'search_terms=referendum' -d "ad_reached_countries=['IE']" \ 1085 -d 'ad_delivery_date_min=2026-01-01' -d 'ad_delivery_date_max=2026-03-31' \ 1086 -d 'fields=page_name,spend,ad_snapshot_url' 1087 1088 # reach, not impressions: the EU and Brazil figures are per-ad totals 1089 curl -s -G "$G" -d "access_token=$TOKEN" \ 1090 -d 'search_page_ids=123456789' -d "ad_reached_countries=['IE']" \ 1091 -d 'fields=eu_total_reach,age_country_gender_reach_breakdown,estimated_audience_size' 1092 1093 # page the result set with the cursor Graph hands back 1094 curl -s -G "$G" -d "access_token=$TOKEN" -d 'search_terms=climate' \ 1095 -d "ad_reached_countries=['GB']" -d 'fields=page_name' -d 'limit=100' -d 'after=CURSOR' 1096 ``` 1097 1098 Leaving out both `search_terms` and `search_page_ids` is a parameter error, not an unfiltered 1099 dump. `search_terms` matches the ad's own language and is capped at 100 characters, spaces are 1100 treated as AND unless you set `search_type=KEYWORD_EXACT_PHRASE`, and `search_page_ids` takes at 1101 most ten IDs. `spend` and `impressions` come back as ranges, not numbers — report them as ranges, 1102 and note that `eu_total_reach` is a single figure while `estimated_audience_size` is another range. 1103 Several fields and filters — `bylines`, `delivery_by_region`, `demographic_distribution`, 1104 `estimated_audience_size_min`/`_max` — exist only for political and issue ads, so an empty result 1105 for a commercial advertiser means the field does not apply rather than that the ad does not exist. 1106 Coverage is political and issue ads in the countries where Meta is required to disclose them, plus 1107 the `EMPLOYMENT_ADS`, `FINANCIAL_PRODUCTS_AND_SERVICES_ADS` and `HOUSING_ADS` categories; 1108 ordinary commercial ads appear in the web Ad Library with far less metadata and are not in the API 1109 the same way. 1110 1111 ### X / Twitter advanced search 1112 1113 Web only in practice. The API is priced out of reach for research, `snscrape` has been 1114 non-functional against X since the 2023 access changes, and the public Nitter instances are 1115 largely dead — so the live tool is X's own advanced search, which still works while logged in. 1116 1117 ```text 1118 from:username since:2026-01-01 until:2026-06-30 one account, bounded by date 1119 to:username replies directed at an account 1120 "exact phrase" -filter:retweets the original post rather than its echoes 1121 filter:images filter:links posts carrying media or outbound links 1122 filter:replies only replies, which is where arguments live 1123 url:example-news-daily.com who linked a domain 1124 geocode:51.5074,-0.1278,5km posts placed near a coordinate 1125 min_faves:500 min_retweets:100 min_replies:50 the versions that actually travelled 1126 lang:ru "phrase" one language, for a multilingual campaign 1127 (term1 OR term2) -term3 grouping and negation, as you would expect 1128 list:12345678 "phrase" search inside a list you have built 1129 conversation_id:1234567890123456789 one thread, including replies 1130 from:username filter:media since:2026-03-01 one account's images in one month 1131 ``` 1132 1133 Every one of these requires a logged-in session, which attributes the searching to that account — 1134 use a research account, and assume the queries are logged. Date-bounded `from:` queries are the 1135 most reliable form; free-text search silently drops older results, so an empty result for an old 1136 date range is not evidence of absence. X publishes no machine-readable operator list and its help 1137 pages block automated fetches, so verify any operator against the advanced-search form itself 1138 before relying on it — the form is the only authority, and it has quietly lost operators before. 1139 For deleted posts, go to the [Wayback Machine](https://web.archive.org/) and archive.today first: 1140 for X specifically, archived snapshots are now a more dependable source than the live site. A 1141 status ID still dates itself with no session at all, as shown under [X / Twitter](#x--twitter) 1142 above. 1143 1144 ## Platform tooling at a glance 1145 1146 | Tool | Platform | Account needed | What it gets you | 1147 | --- | --- | --- | --- | 1148 | [yt-dlp](https://github.com/yt-dlp/yt-dlp) | 1,000+ sites | No, until the platform insists | Video and channel metadata, subtitles, comments, timed clips | 1149 | [gallery-dl](https://github.com/mikf/gallery-dl) | Images and galleries, most platforms | No, until the platform insists | Bulk profile capture with per-item JSON sidecars | 1150 | [Instaloader](https://instaloader.github.io/) | Instagram | No for public posts, yes for stories | Posts, captions, comments, geotags, profile pictures | 1151 | [instagram-location-search](https://github.com/bellingcat/instagram-location-search) | Instagram | Yes, a session cookie | Location IDs near a coordinate, as CSV, JSON, GeoJSON or a map | 1152 | [Telethon](https://docs.telethon.dev/) | Telegram | Yes, API credentials and a phone number | Message history, forward graph, member lists, creation dates | 1153 | [Telepathy](https://github.com/proseltd/Telepathy-Community) | Telegram | Yes, API credentials and a phone number | The same as a CLI, but unmaintained since July 2024 | 1154 | [Telegram Phone Number Checker](https://github.com/bellingcat/telegram-phone-number-checker) | Telegram | Yes, and it writes to your contacts | Whether a number or username is registered | 1155 | [tiktok-hashtag-analysis](https://github.com/bellingcat/tiktok-hashtag-analysis) | TikTok | No | Posts by hashtag, plus a co-occurrence table or plot | 1156 | [Arctic Shift](https://arctic-shift.photon-reddit.com/) | Reddit | No | Deleted posts and comments, wiki page lists, activity aggregates | 1157 | [goat](https://github.com/bluesky-social/goat) | Bluesky / AT Protocol | No for reads | Repository and blob exports, PLC history, TID decoding, firehose | 1158 | [Meta Ad Library API](https://developers.facebook.com/docs/graph-api/reference/ads_archive/) | Facebook, Instagram | Yes, an approved developer app | Spend and reach ranges, payer bylines, delivery breakdowns | 1159 | [X/Twitter Advanced Search](https://x.com/search-advanced) | X | Yes, a logged-in session | Date-, geo- and engagement-bounded post search | 1160 1161 Reach for `yt-dlp` and `gallery-dl` first on any platform either covers, because they are the only 1162 two on this list maintained at the pace platforms break things. Reach for Arctic Shift and `goat` 1163 when the question is historical, because both hold data the live platform will not serve. Everything 1164 else here is worth a check against its last commit date before you trust its output. 1165 1166 ## Tool reference 1167 1168 | Tool | What it does | Cost | 1169 | --- | --- | --- | 1170 | [4plebs](https://4plebs.org/) | Searchable archive of specific 4chan boards. Makes it possible to read threads after they are purged from 4chan. | free | 1171 | [Bellingcat TikTok Date Extract](https://bellingcat.github.io/tiktok-timestamp) | Get the exact upload date + time for tiktok video urls | free | 1172 | [Bellingcat TikTok Hashtag Analysis](https://github.com/bellingcat/tiktok-hashtag-analysis) | Archive content and metadata from TikTok posts that contain one or more specified hashtags | free | 1173 | [Blackbird](https://github.com/p1ngul1n0/blackbird) | Check usernames and email addresses on websites and social networks | free | 1174 | [BskyFollowFinder](https://bsky-follow-finder.theo.io/) | A tool that identifies which Bluesky accounts are followed by a profile’s contacts but not by that profile. Can be used for expanding networks and social… | free | 1175 | [Disboard](https://disboard.org/servers) | Search for public discord servers | free | 1176 | [Discord Chat Exporter](https://github.com/Tyrrrz/DiscordChatExporter) | A tool for exporting and backing up Discord chat logs in multiple formats. | free | 1177 | [DiscordLeaks](https://discordleaks.unicornriot.ninja/) | Search hundreds of thousands of messages leaked from 290+ white-supremacist / nazi discord servers. | free | 1178 | [F5Bot](https://f5bot.com/) | Sends you an email when a keyword is mentioned on Reddit. | free | 1179 | [Facebook Video Downloader](http://fdown.net/) | Handy website to download public Facebook videos. Copy paste the URL of the video and download it in the available definition formats. | free | 1180 | [FindClone](https://findclone.ru/) | Searches images from VK profiles (within certain limits) | free | 1181 | [Ghunt](https://github.com/mxrch/GHunt) | A command line tool for obtaining information about Google accounts. | free | 1182 | [Google Account Finder (EPIEOS)](https://tools.epieos.com/google-account.php) | Find the profile picture and public Google Map Reviews + Photos associated with a G-mail adress. Also checks for phone numbers, and checks for email… | free | 1183 | [Gravatar Email Checker](https://en.gravatar.com/site/check/) | Check if an email address has been used to comment on blogs and whether there is a profile image attached. | free | 1184 | [HaveIBeenZuckered](https://haveibeenzuckered.com/) | Check if a telephone number is present within the Facebook data breach. | free | 1185 | [Hoaxy](https://x.com/OSoMe_IU/status/1949394974262350098) | Hoaxy is a web-based search and visualization tool. It helps visualize the spread of information on Bluesky and X (Twitter). | partly free | 1186 | [Instagram Location Search](https://github.com/bellingcat/instagram-location-search/tree/main) | A command line tool that allows users to find location tags near a specified latitude and longitude. | free | 1187 | [InstaLoader](https://instaloader.github.io) | Download pictures or videos (with metadata) from Instagram. | free | 1188 | [Intelligence X Telegram SearchDDD](https://intelx.io/tools?tab=telegram) | Google-based search engine for Telegram (includes Telegago) | free | 1189 | [LinkdTime](https://github.com/Lucksi/LinkdTime) | Build a clean timeline of any LinkedIn activity from a single URL or a whole list of links. | free | 1190 | [Meta Content Library](https://transparency.meta.com/researchtools/meta-content-library) | Meta Content Library is a controlled-access tool that lets approved academic and non-profit researchers search the full public archive of Facebook… | free | 1191 | [MW Geofind](https://mattw.io/youtube-geofind/location) | MW Geofind is a tool designed to help users identify the filming location of YouTube videos, facilitating the exploration of global content from a… | free | 1192 | [Open Measures](https://public.openmeasures.io) | Open Measures helps open source researchers investigate harmful online activity such as extremism and disinformation. | partly free | 1193 | [Photo-Map.RU](http://photo-map.ru/) | Geotagged VK posts. | free | 1194 | [PSNprofiles](https://psnprofiles.com/) | Search PlayStation username, see daily activity, games played, country, and profile pic | free | 1195 | [RadiTube](https://tool.raditube.com/) | A search engine that searches the subtitles of about 380 (right/left) radical YouTube channels. You query for example for "q says" of "voter fraud" in… | free | 1196 | [RedditMetis](https://redditmetis.com) | RedditMetis is an online tool that analyses any public Reddit profile and returns charts and statistics on the user’s activity, language, sentiment and… | free | 1197 | [Sherlock](https://github.com/sherlock-project/sherlock) | Allows a user to search for the presence of specific usernames across more than 400 websites and social networks. | free | 1198 | [Snap Map](https://map.snapchat.com) | Searchable map of geotagged snaps. | free | 1199 | [Social Searcher](https://www.social-searcher.com/) | Search hashtags and usernames across various platforms. | partly free | 1200 | [SteamId.uk](http://steamid.uk/) | Lookup player names, view (more) previously used names, and when accounts befriended eachother (Free). View screenshots of account, (bulk) seach based on… | partly free | 1201 | [Story Saver](https://storysaver.net) | Download public Instagram Stories, Highlights and Videos. | free | 1202 | [Strava](https://www.strava.com) | A fitness tracking platform where publicly shared GPS activity data can reveal movement patterns, routines, and precise locations of individuals… | partly free | 1203 | [Telegago](https://cse.google.com/cse?cx=006368593537057042503:efxu7xprihg) | Telegago is a Google Custom Search Engine tailored for searching public Telegram content for OSINT purposes. | free | 1204 | [Telegram Group Joiner](https://bellingcat.github.io/telegram-group-joiner/) | Automate joining multiple Telegram groups and channels, ideal for researchers monitoring specific topics. | free | 1205 | [Telegram Phone Number Checker](https://github.com/bellingcat/telegram-phone-number-checker) | Command line tool for checking if phone numbers are connected to Telegram accounts and retrieving related information where available. | free | 1206 | [TelegramDB](https://www.telegramdb.org/) | TelegramDB is a searchable database service that allows users to explore public Telegram groups and channels via a dedicated bot. | partly free | 1207 | [Telemetrio](https://telemetr.io/) | Telemetr.io offers a range of Telegram-related services based on a catalog of Telegram channels: country and category-specific rankings, curated… | partly free | 1208 | [Telemetry](https://www.telemetryapp.io/) | An analytical search tool for Telegram groups and channels. | partly free | 1209 | [Telepathy](https://github.com/prose-intelligence-ltd/Telepathy-Community) | Telepathy is a versatile Telegram toolkit for OSINT analysts, enabling chat archiving, memberlist gathering, user location lookup, top poster analysis… | free | 1210 | [TGStat](https://tgstat.com/) | TGStat is a web-based analytics tool for Telegram that monitors active channels and provides profile analytics and statistics. It tracks channel… | partly free | 1211 | [TikTok Ad Library](https://library.tiktok.com/ads) | Search and review TikTok ads and related metadata for transparency and research. | free | 1212 | [TikTok-Api](https://github.com/davidteather/TikTok-Api/) | TikTok-Api is an open-source Python library (unofficial TikTok API) that allows developers and researchers to retrieve various public data from TikTok | free | 1213 | [tlgrm.eu channels](http://tlgrm.eu/channels) | Search Telegram channels. | free | 1214 | [Twitter Video Downloader](https://twittervideodownloader.com/) | Download videos from X (formerly Twitter) by converting tweet URLs into downloadable video links. | free | 1215 | [Twitter/X Location Search](https://twitter.com/explore) | Search for geocoded tweets by their distance from some coordinates. | free | 1216 | [Vk.watch](http://vk.watch/) | See public comments left by an account, profile photos used, and very basic facial recognition | free | 1217 | [Who posted what?](https://whopostedwhat.com/) | A tool that allows a keyword search on Facebook on a specific date or within a specific time frame. | free | 1218 | [X/Twitter Advanced Search](https://x.com/search-advanced) | Twitter/X Advanced Search is X's own tool to help users find more precise information on the platform by filtering posts according to criteria such as… | free | 1219 | [XboxGamertag](https://xboxgamertag.com/) | Search gamertags, see games played and recorded game clips | free | 1220 | [YouTube Metadata](https://mattw.io/youtube-metadata/) | An alternative to Amnesty's YT viewer, with slightly more information. | free | 1221 1222 ## Pitfalls 1223 1224 - **Viewing can notify.** LinkedIn tells people who looked. Story views are attributed. Joining a 1225 Telegram group adds you to a visible list. 1226 - **Tooling rots fast.** Anything depending on an undocumented endpoint is one deploy from broken. 1227 Check a tool's commit history before trusting its output. 1228 - **Deleted is not gone, and present is not original.** Archive first, then analyse. 1229 - **Reposts dominate.** The account with the most engagement is rarely the origin. Trace back. 1230 - **An ID dates the record, not the footage.** Decoding a TikTok or Snowflake ID tells you when 1231 that upload was created. A 2019 clip uploaded in 2026 has a 2026 ID and the arithmetic is still 1232 right. 1233 - **A client-written timestamp is a claim.** An atproto TID and `record.createdAt` are both chosen 1234 by whatever wrote the record; `indexedAt` is the relay's. Quote the one the account did not 1235 control, and say which you are quoting. 1236 - **Zero results and broken scraper look identical.** A busy hashtag returning no posts, or a 1237 `--match-filters` expression naming a field that extractor does not emit, both print nothing. 1238 Run one query you already know the answer to before you believe a negative. 1239 - **A pipeline can match and then discard the match.** `grep -o '[0-9]*$'` against 1240 `data-post="channel/441"` anchors digits to a line that ends in a quote, so it emits a blank 1241 line per post and exits 0. The extractor worked, the shell threw the answer away, and the 1242 symptom is an empty file. Count the lines at each stage of a new pipeline before trusting the 1243 last one. 1244 - **A date filter may not filter on the date you mean.** Bluesky's `since`/`until` on 1245 `searchPosts` are documented against the server's `sortAt`, which need not equal 1246 `record.createdAt`; Arctic Shift's `created_utc` aggregate buckets are cut in a DST-observing 1247 local zone. Both return a defensible answer to a question you did not ask. 1248 - **Unauthenticated does not mean unattributed.** Your IP is an identifier, and a sweep from one 1249 address is a signature. Conversely, authenticating to get past a block converts an anonymous read 1250 into a logged one tied to an account you will lose. 1251 - **Datacentre addresses are blocked where browsers are not.** Reddit's `.json` surface returns 403 1252 from hosted ranges regardless of User-Agent. Code that works on a laptop and fails in CI has not 1253 found a bug; it has found a policy. 1254 - **Archives ingest, they do not watch.** Arctic Shift holds what it captured when it captured it: 1255 an edit made afterwards is invisible, and something deleted within seconds may never have been 1256 recorded at all. 1257 - **Spend and reach are ranges.** Meta reports `spend` and `impressions` as buckets. Writing a 1258 midpoint and calling it a figure invents precision the source does not have. 1259 - **Counts from a sample are comparative at best.** Hashtag collection, search results and "top 1260 twenty" lists are samples whose size you do not control. Two runs an hour apart will disagree, 1261 which is a property of the method, not of the world. 1262 - **The same handle on four platforms is one lead, not one person.** It takes content, timing or a 1263 payment record to tie accounts together. Say which of the three you have. 1264 1265 ## Worked example 1266 1267 One datum: the handle **@civicwatchnow**, read off a screenshot of a post, with no platform 1268 attached and no other context. 1269 1270 1. **Find the platforms.** A username sweep — see 1271 [username and account discovery](/sheets/osint/usernames-and-accounts) — returns hits on 1272 Bluesky, YouTube, Reddit and a Telegram channel of the same name. Four surfaces, four different 1273 amounts of openness. 1274 2. **Date the accounts, keylessly.** `resolveHandle` gives the Bluesky DID 1275 `did:plc:z72i7hdynmk6r22z27h6tvur`; the PLC audit log's first entry dates the account to 1276 **12 March 2026** and shows one earlier handle. An account presenting itself as an established 1277 watchdog, three months old, is the first real finding. 1278 3. **Characterise the output.** `yt-dlp --flat-playlist --dump-json` on the YouTube channel lists 1279 **41 videos**, all uploaded between 2026-04-02 and 2026-05-14 — six weeks — with `upload_date` 1280 clustered on weekdays. Volume and cadence, from the platform's own metadata. 1281 4. **Read what was deleted.** Arctic Shift's `comments/search?author=civicwatchnow&limit=100` 1282 returns **300 comments across eleven subreddits**, of which fourteen no longer exist on Reddit. 1283 The removed ones carry the same link, which the live profile does not show. The aggregate 1284 endpoint gives the shape of it: 1285 1286 ```json 1287 {"data":[{"created_utc":"2026-03-31T22:00:00.000Z","count":"18"}, 1288 {"created_utc":"2026-04-30T22:00:00.000Z","count":"197"}, 1289 {"created_utc":"2026-05-31T22:00:00.000Z","count":"85"}]} 1290 ``` 1291 1292 5. **Map the amplification.** The Telethon forward counter over the Telegram channel's last 2,000 1293 messages returns a top-twenty list in which **three channels account for 1,340 of 1,612 1294 forwards**, and the channel's own `entity.date` is **2026-03-14**, two days after the Bluesky 1295 DID's first PLC operation. 1296 6. **Check the gaps.** `page.total` reports 1,900 messages against a newest ID of **2,146**, so 1297 roughly 246 messages — 11 per cent — have been deleted. Which ones is a separate question; that 1298 they were is not. 1299 7. **Check for paid distribution.** The Meta Ad Library `search_terms=civicwatch` with 1300 `ad_reached_countries=['GB']` and `ad_active_status=ALL` returns two advertisers running the 1301 same creative, with `bylines` naming a company and `spend` as the range **£5,000–£9,999**. That 1302 is a funding lead the organic surfaces never gave. 1303 8. **Trace to origin, then archive.** The earliest instance of the clip is the TikTok upload, dated 1304 from the video ID — `7169068344567909638 >> 32` = **2022-11-23 04:46:37Z** — rather than the 1305 display text, which read "2 days ago". The YouTube and Telegram copies are later. Capture all 1306 four profiles, `goat repo export` the Bluesky repository and save the ad snapshots before 1307 writing anything — see [archiving and evidence](/sheets/osint/archiving-and-evidence). 1308 1309 What you can assert: four accounts under one handle, created within three days of each other in 1310 March 2026, pushing a single clip whose earliest known copy is dated on TikTok to November 2022, 1311 amplified by three Telegram channels carrying 83 per cent of the forwards, and promoted by two 1312 named advertisers with a disclosed spend range. Every date in that sentence came from a 1313 platform-side identifier or an archive, not from a rendered page. 1314 1315 What would falsify it: an earlier copy of the clip on a platform nobody swept, which would make the 1316 TikTok ID a reupload's ID rather than the original's; a Telegram channel that was renamed into the 1317 handle rather than created under it, which the creation date alone cannot distinguish; or an 1318 advertiser byline that turns out to be a media-buying agency acting for an unrelated client. The 1319 measurement carrying the most risk is the Bluesky creation date, because a PLC log's first 1320 operation dates the *DID*, not the project — an operator who registered the DID months before 1321 using it would push that date earlier with no deception at all, so quote it as "first PLC operation 1322 on 12 March 2026" rather than as a founding date. What you cannot assert at all is that one person 1323 runs all four: a shared handle is a lead until content, timing or a payment record ties them, and 1324 here it is the ad byline, not the handle, doing that work. 1325 1326 ## Broader catalogues 1327 1328 - [Social Media OSINT](https://tools.osintnewsletter.com/tool-categories/social-media-osint) 1329 1330 1331 ## More tools 1332 1333 Further tools for this area from the OSINT Newsletter Tools Library ([Facebook](https://tools.osintnewsletter.com/tool-categories/social-media-osint/facebook), [Instagram](https://tools.osintnewsletter.com/tool-categories/social-media-osint/instagram), [Twitter/X](https://tools.osintnewsletter.com/tool-categories/social-media-osint/twitter-x), [Telegram](https://tools.osintnewsletter.com/tool-categories/social-media-osint/telegram), [YouTube](https://tools.osintnewsletter.com/tool-categories/social-media-osint/youtube), [Reddit](https://tools.osintnewsletter.com/tool-categories/social-media-osint/reddit), [Bluesky](https://tools.osintnewsletter.com/tool-categories/social-media-osint/bluesky), [Snapchat](https://tools.osintnewsletter.com/tool-categories/social-media-osint/snapchat), [Other/Multiple Platforms](https://tools.osintnewsletter.com/tool-categories/social-media-osint/other-multiple-platforms)), excluding those already listed above. 1334 1335 | Tool | What it does | 1336 | --- | --- | 1337 | [Analyst Research Tools](https://analystresearchtools.com/) | A browser-based investigative research platform pulling together a collection of free lookup tools to help you gather, organise… | 1338 | [Archivarix Tube Search](https://tube.archivarix.net/) | A free OSINT tool for discovering historical versions of YouTube videos using archived snapshots from archive data. | 1339 | [Arctic Shift](https://arctic-shift.photon-reddit.com/) | A Reddit search and archival platform that enables investigators to search historical Reddit posts, comments, deleted content… | 1340 | [Aware Online: Reddit Search](https://www.aware-online.com/en/osint-tools/reddit-search-tool/) | A free OSINT tool that helps investigators search Reddit more effectively by querying posts, comments, usernames, and discussions… | 1341 | [AzSky](https://azsky.app/) | A Reddit-style interface powered by Bluesky, designed to make exploring decentralised social conversations easier. | 1342 | [BirdHunt](https://birdhunt.huntintel.io/) | Find tweets based on location, helping you uncover what people were saying in a specific place at a specific time. | 1343 | [Bluesky Follow Finder](https://bsky-follow-finder.theo.io/) | Surfaces accounts followed by people a profile follows, but that the profile doesn’t. | 1344 | [BoardReader](https://boardreader.com/) | A specialised search engine for online forums, message boards and discussion communities. | 1345 | [Buzzglobe](https://buzzglobe.com/) | Social-media search aggregator for discovering publicly indexed profiles, pages, posts and images across multiple platforms. | 1346 | [de-digger](https://www.dedigger.com/#gsc.tab=0) | A specialist search and discovery engine for finding publicly accessible files on Google Drive. | 1347 | [Deaddrop](https://deaddrop.theosintconsultants.com/login) | Telegram search engine. | 1348 | [Discord Leaks](https://discordleaks.unicornriot.ninja/) | A searchable archive of leaked Discord server, RocketChat and Skype messages, mainly from extremist and politically active… | 1349 | [Dolphin Radar](https://www.dolphinradar.com/) | All-in-one Instagram Tracker & Analyser | 1350 | [Dumpor](https://dumpor.io/) | Instagram viewer and profile analysis tool that enables anonymous viewing of public Instagram content without logging in. | 1351 | [Export Comments](https://exportcomments.com/) | Extracts, downloads, and analyses comments from social media platforms and video-sharing sites. | 1352 | [ExportGram](https://exportgram.net/) | A web tool that lets you extract and download comments from Instagram posts and reels into spreadsheet formats. | 1353 | [Filmot](https://www.reddit.com/r/languagelearning/comments/odj2gx/ive_built_a_search_engine_across_youtube_captions/) | A search engine that lets you search through YouTube video subtitles and metadata to find videos that contain specific words or… | 1354 | [Govsky](https://www.govsky.org/) | Allows users to explore and identify official government and public sector accounts on the Bluesky social media platform. | 1355 | [Graph Tips](https://graph.tips/) | Generates Facebook search queries that recreate much of the functionality of the former Facebook Graph Search. | 1356 | [Hippie OSINT Toolkit](https://osint.hippie.cat/) | An OSINT web toolkit that allows you to reverse search a domain, a TikTok post, an image, or username (and more). | 1357 | [Inflact Instagram Profile Viewer](https://inflact.com/instagram-viewer/profile/) | A no-login Instagram viewer that lets you peek at public profiles, posts, and basic engagement data anonymously. | 1358 | [Instagram Map](https://www.youtube.com/watch?v=R_oOdBTBkGg) | Allows users within Instagram to view geotagged posts, Stories, and shared live locations (where enabled) on a map inside the… | 1359 | [Internect.info](https://internect.info/) | A search interface for the AT Protocol allowing users to locate and query Bluesky handles and associated account information… | 1360 | [Nitter](https://nitter.net/) | A privacy-focused alternative front-end for X (Twitter) that lets you view tweets without tracking, ads, or needing an account. | 1361 | [Sotwe](https://www.sotwe.com/) | Tracks Twitter trends and finds the most popular Twitter users, hashtags and places. | 1362 | [Sowsearch](https://www.sowsearch.info/) | A specialist Facebook tool to search posts, profiles, comments, places, events, groups, and public content more efficiently using… | 1363 | [Story Saver (Instagram)](https://www.storysaver.net/en/) | Download public Instagram Stories, highlights and videos. | 1364 | [Subreddit Stats](https://subredditstats.com/) | A free Reddit analytics tool used to track subreddit growth, activity levels, top posts/users, keyword trends, and community… | 1365 | [Subreddits.org](https://subreddits.org/) | An online directory and search tool that helps users find and explore Reddit communities (subreddits) by keyword, topic, or… | 1366 | [Telegago (Telegram)](https://cse.google.com/cse?cx=006368593537057042503:efxu7xprihg#gsc.tab=0) | A custom Google search engine to help users search publicly available content on Telegram channels & groups. | 1367 | [Telegram Spoiler Decoder](https://spoiler.soxoj.com/) | Instantly recovers hidden spoiler text on MacOS within a Telegram chat. | 1368 | [ToolsCord](https://toolscord.com/) | A free, browser-based toolkit focused on Discord utilities. | 1369 | [Treeverse](https://treeverse.app/) | Visualise and explore conversations on Bluesky as interactive threads and networks. | 1370 | [TweeterID](https://tweeterid.com/) | A simple X/Twitter username and ID conversion tool. | 1371 | [Twitter Viewer](https://twitterwebviewer.com/) | An online tool that lets you browse Twitter profiles, tweets, and media, and download videos without logging in. | 1372 | [Waybien](https://waybien.com/en) | A search engine for Telegram, Facebook, Discord, and Whatsapp groups. | 1373 | [Who Posted What? Facebook Tool](https://whopostedwhat.com/) | A non-public Facebook keyword search tool that allows you to search keywords on specific dates. | 1374 | [YouTube Video Finder](https://findyoutubevideo.thetechrobo.ca/) | Investigates a YouTube video using its URL or ID, with quick links to archives and analysis tools. | 1375 1376 ## Sources 1377 1378 Both catalogues below are maintained by other people and are considerably larger than 1379 this page. Use them as the canonical index; this sheet is a working route through them. 1380 1381 - [Bellingcat's Online Investigation Toolkit](https://bellingcat.gitbook.io/toolkit) — ~340 tools, each with its own 1382 review page covering cost, difficulty, requirements and limitations. 1383 - [OSINT Newsletter Tools Library](https://tools.osintnewsletter.com) — ~280 tools, organised by investigative goal. 1384 1385 Neither publishes a licence, so nothing here is copied from them: tool names, one-line 1386 descriptions, cost flags and links are catalogue facts, and the method and commentary are 1387 this site's own. See [credits](/credits).