daemon-sec-cheatsheet

The cheatsheet vault for operators: AD, enumeration, exploitation, priv-esc, web, DFIR
git clone https://git.daemon-sec.xyz/daemon-sec-cheatsheet.git
Log | Files | Refs | README | LICENSE

social-media-platforms.md (85401B)


      1 ---
      2 title: "Social Media — Per-Platform Research"
      3 description: "What each major platform still exposes to an unauthenticated researcher, and the tools that work per platform."
      4 category: osint
      5 subcategory: "Social Media"
      6 tags: [osint, social-media, telegram, tiktok]
      7 tools: [yt-dlp, gallery-dl, instaloader, telethon, telepathy, arctic-shift, atproto, goat, tiktok-hashtag-analysis, instagram-location-search]
      8 difficulty: intermediate
      9 updated: 2026-10-04
     10 references:
     11   - name: "Bellingcat's Online Investigation Toolkit"
     12     url: "https://bellingcat.gitbook.io/toolkit"
     13     author: "Bellingcat"
     14     license: none
     15     relation: derived
     16     note: "Tool catalogue: names, descriptions, cost flags and links for this area."
     17   - name: "OSINT Newsletter Tools Library"
     18     url: "https://tools.osintnewsletter.com"
     19     author: "The OSINT Newsletter"
     20     license: none
     21     relation: derived
     22     note: "Second tool catalogue, cross-checked against the above."
     23   - name: "yt-dlp option reference"
     24     url: "https://github.com/yt-dlp/yt-dlp#usage-and-options"
     25     author: "yt-dlp contributors"
     26     relation: link-only
     27     note: "Every yt-dlp flag on this page was checked against the project's own option list."
     28   - name: "gallery-dl command-line options"
     29     url: "https://github.com/mikf/gallery-dl/blob/master/docs/options.md"
     30     author: "Mike Fährmann"
     31     relation: link-only
     32     note: "Every gallery-dl flag on this page was checked against the project's own option list."
     33   - name: "Instaloader command-line options"
     34     url: "https://instaloader.github.io/cli-options.html"
     35     author: "Alexander Graf, André Koch-Kramer"
     36     relation: link-only
     37     note: "Instaloader flags and target syntax checked against the project's own documentation."
     38   - name: "Arctic Shift API reference"
     39     url: "https://github.com/ArthurHeitmann/arctic_shift/blob/master/api/README.md"
     40     author: "Arthur Heitmann"
     41     relation: link-only
     42     note: "Endpoint paths and query parameters checked against the project's own API reference."
     43   - name: "AT Protocol specifications and lexicons"
     44     url: "https://atproto.com/specs/tid"
     45     author: "Bluesky Social"
     46     relation: link-only
     47     note: "XRPC parameter names and the TID timestamp layout checked against the published lexicons."
     48   - name: "Meta Ad Library API reference"
     49     url: "https://developers.facebook.com/docs/graph-api/reference/ads_archive/"
     50     author: "Meta Platforms"
     51     relation: link-only
     52     note: "ads_archive parameters, enum values and field names checked against the API reference."
     53 ---
     54 
     55 ## What this covers
     56 
     57 Each platform exposes a different surface and closes it on a different schedule. This is the
     58 per-platform picture: what is still reachable without an account, what needs one, and which tools
     59 survive contact with the current APIs. Assume anything here can break without notice — platforms
     60 change access rules faster than tooling keeps up.
     61 
     62 ## Method
     63 
     64 1. **Decide whether you can be seen before you touch anything.** Some platforms attribute a view,
     65    a follow or a group join to your account. Work out which, on this platform, today — then pick
     66    the account you are willing to burn.
     67 2. **Archive first, analyse second.** Capture the post, the profile and the surrounding thread
     68    before you start pulling threads. Accounts get deleted mid-investigation, routinely.
     69 3. **Take the public surface without an account** where it exists: profile metadata, follower
     70    counts, pinned posts, video metadata. Much of this is still reachable unauthenticated and
     71    costs you nothing.
     72 4. **Pivot on identifiers, not display names.** Numeric IDs, DIDs and channel IDs survive renames
     73    and vanity-URL changes; handles and names do not.
     74 5. **Trace to origin.** The account with the most engagement is almost never the source. Follow
     75    forwards, quote-posts and reuploads backwards until the chain stops.
     76 6. **Date everything from the platform, not the page.** Upload timestamps, message IDs and
     77    creation dates are platform-side facts; the date shown on a reposted screenshot is not.
     78 7. **Expect the tooling to be broken.** Anything built on an undocumented endpoint dies without
     79    notice. Check the project's last commit before you trust its output, and cross-check one result
     80    by hand.
     81 
     82 The judgement calls: decide early whether you need the content or the network, because bulk media
     83 capture and graph enumeration use different tools and different amounts of patience. Decide which
     84 account you are prepared to attribute the work to, before you authenticate anything. And decide
     85 whether a date is given or derived — if you read a date off the page and then "confirm" it with a
     86 tool that reads the same page, you have proved nothing.
     87 
     88 ## The identifier arithmetic
     89 
     90 Most platform IDs are not counters. They are clock readings, so the ID you already have dates the
     91 object with no request to the platform at all. That matters twice over: the object may be deleted
     92 by the time you look, and the date rendered on a page is frequently the repost's date rather than
     93 the original's.
     94 
     95 **TikTok.** A video ID is a 64-bit value whose top 32 bits hold a Unix timestamp in seconds. Shift
     96 right by 32 and read it:
     97 
     98 ```text
     99 video id     : 7169068344567909638
    100 width        : 63 bits, so the timestamp occupies bits 62..32
    101 id >> 32     : 1669178797
    102 as UTC       : 2022-11-23 04:46:37Z
    103 ```
    104 
    105 **X / Twitter.** Status IDs are Snowflakes: 41 bits of milliseconds since a platform-specific
    106 epoch, sitting 22 bits up. The epoch is the number you have to get right.
    107 
    108 ```text
    109 status id          : 1949394974262350098
    110 id >> 22           : 464771979871           milliseconds since the Twitter epoch
    111 + 1288834974657    : 1753606954528          milliseconds since the Unix epoch
    112 as UTC             : 2025-07-27 09:02:34.528Z
    113 ```
    114 
    115 **Bluesky.** A record key is a TID: 13 characters of the sortable base32 alphabet
    116 `234567abcdefghijklmnopqrstuvwxyz`, decoding to a top bit of 0, then 53 bits of **microseconds**
    117 since the Unix epoch, then a 10-bit random clock identifier. Because the alphabet sorts, record
    118 keys sort chronologically as plain strings.
    119 
    120 ```bash
    121 # all three, from the shell
    122 python3 -c 'import datetime as d; i=7169068344567909638; print(d.datetime.fromtimestamp(i>>32, d.timezone.utc))'
    123 python3 -c 'import datetime as d; i=1949394974262350098; print(d.datetime.fromtimestamp(((i>>22)+1288834974657)/1000, d.timezone.utc))'
    124 goat syntax tid inspect 3kzifvcppte22
    125 ```
    126 
    127 ```text
    128 2022-11-23 04:46:37+00:00
    129 2025-07-27 09:02:34.528000+00:00
    130 Timestamp (UTC): 2024-08-12T02:08:03.29Z
    131 Timestamp (Local): 2024-08-11T19:08:03-07:00
    132 ClockID: 0
    133 uint64: 0x187dcbda2b5ca800
    134 ```
    135 
    136 Three ways this misleads. The first is the obvious one: an ID dates the **record**, not the
    137 footage. A clip shot in 2019 and uploaded in 2026 has a 2026 ID, and the arithmetic is still
    138 correct. The second is who generated the number. TikTok and X mint IDs server-side at creation, so
    139 they are harder to forge than a rendered date; an atproto TID is minted by whichever component
    140 writes the record, off that generator's own local clock, and the specification states plainly that
    141 uniqueness cannot be guaranteed because an adversarial party can create records reusing known TIDs.
    142 A Bluesky record key is therefore the same class of evidence as `record.createdAt` — an assertion
    143 by the writer — and the relay's `indexedAt` is the independent number. The third is bit-width drift: quote the shift, not a
    144 string slice. Taking "the first 31 characters of the binary" works for every current 19-digit
    145 TikTok ID only because such IDs happen to be 63 bits wide; `id >> 32` keeps working when they are
    146 not.
    147 
    148 ## Telegram
    149 
    150 The most productive platform for open-source research, because public channels are genuinely
    151 public, history is retained, and forwarding metadata exposes the network between channels.
    152 
    153 - **Forward chains** are the key primitive. A forwarded message keeps its origin, so you can trace
    154   content back through the channels that amplified it.
    155 - **Member lists** are visible in many groups. Your own account appears in them too — use a
    156   research account.
    157 - **Channel creation dates and message IDs** are sequential, which lets you estimate when a channel
    158   started and how much has been deleted.
    159 
    160 Before any tooling, the web preview gives you a channel's recent history with no account, no app
    161 and no API credentials. It is the cheapest first look there is, and it is the only one that costs
    162 you nothing in attribution:
    163 
    164 ```text
    165 https://t.me/s/channelname              the public web view: recent messages, dates, view counts
    166 https://t.me/channelname/4312           one message, by its sequential id -- the permalink form
    167 https://t.me/s/channelname?before=4312  page backwards from that id, ~20 messages at a time
    168 https://t.me/+AbCdEf0123456789          a private invite link; opening it does not join, but
    169                                         it does reveal the chat title and member count
    170 ```
    171 
    172 ```bash
    173 # how far back the web view goes, and whether ids are contiguous
    174 curl -s 'https://t.me/s/channelname' \
    175   | grep -o 'data-post="channelname/[0-9]*"' | sed 's/.*\///; s/"//' | sort -n | head -1
    176 
    177 # walk backwards in pages of ~20 and collect every id the preview will serve
    178 for b in 4312 4292 4272 4252; do
    179   curl -s "https://t.me/s/channelname?before=$b" \
    180     | grep -o 'data-post="channelname/[0-9]*"' | sed 's/.*\///; s/"//'
    181 done | sort -nu > ids.txt
    182 
    183 # the gaps are deletions: compare the id range against the count you actually got
    184 awk 'NR==1{min=$1} {max=$1; n++} END{print "ids", min"-"max, "present", n, "missing", max-min+1-n}' ids.txt
    185 ```
    186 
    187 ```text
    188 ids 401-440 present 39 missing 1
    189 ```
    190 
    191 Strip the id with `sed`, not with a second `grep -o '[0-9]*$'`. The attribute ends in a quote, so
    192 anchoring digits to end-of-line matches the empty string and emits one blank line per post: twenty
    193 matches in, twenty blanks out, exit status 0. The run looks like a channel with no posts rather
    194 than like a pipeline that threw the answer away.
    195 
    196 The web view is a *preview*, not the archive: it serves a bounded window of recent history and
    197 will not page back indefinitely, and a channel can disable it entirely. Treat a missing ID as
    198 evidence of deletion only after you have confirmed the preview is serving that range at all.
    199 Archiving a channel properly, mapping its forwards and checking a number against the platform all
    200 have tooling below under [Key tools](#key-tools).
    201 
    202 ## Facebook
    203 
    204 Graph search is gone and most enumeration is dead. What still works:
    205 
    206 - **Numeric profile IDs** persist even when vanity URLs change, and map back to a profile.
    207 - **Public pages, groups and events** remain readable and are frequently forgotten by their owners.
    208 - **Ad Library** ([facebook.com/ads/library](https://www.facebook.com/ads/library/)) is genuinely
    209   open, keyword-searchable, and covers political advertising with spend and reach data.
    210 - **Photo comments and reactions** on public posts still expose participant lists.
    211 
    212 The two URL forms worth memorising. The first resolves a numeric ID that survived a rename; the
    213 second is the Ad Library filter set, which the web interface writes into its own query string, so
    214 you can build a link rather than click through four dropdowns:
    215 
    216 ```text
    217 # a numeric id back to a profile, which survives every vanity-URL change
    218 https://www.facebook.com/profile.php?id=100012345678901
    219 
    220 # the Ad Library UI's own filter parameters, as it constructs them
    221 https://www.facebook.com/ads/library/?active_status=all&ad_type=political_and_issue_ads
    222     &country=GB&q=climate&search_type=keyword_unordered&media_type=all
    223 
    224 # every ad from one page, which is the attribution-relevant view
    225 https://www.facebook.com/ads/library/?active_status=all&ad_type=all&country=GB
    226     &view_all_page_id=123456789
    227 ```
    228 
    229 Those are interface parameters observed from the Ad Library itself, not a documented API; the
    230 documented parameter set is the Graph endpoint under
    231 [Meta Ad Library API](#meta-ad-library-api) below, and that is what to script against. Everything
    232 else on Facebook that once had a query form now does not: [Who posted what?](https://whopostedwhat.com/)
    233 and [Graph Tips](https://graph.tips/) build search URLs for you, and whether they work on any given
    234 day is a question you answer by trying.
    235 
    236 ## Instagram
    237 
    238 - Profile pictures, bios and post counts are visible without login; the feed usually is not.
    239 - `?__a=1` style JSON endpoints have been closed off repeatedly — assume they do not work.
    240 - [Instaloader](https://instaloader.github.io/) still handles public profile and hashtag
    241   downloads, including comments and geotags where present; it has its own section below.
    242 - Aggressive use gets the account and the IP rate-limited quickly, then blocked outright.
    243 
    244 Instaloader's target syntax is the part worth learning, because it is how you address something
    245 other than a profile. All four forms are documented targets, not flags:
    246 
    247 ```bash
    248 profile_name         a public profile, or a private one with --login
    249 "#hashtag"           posts carrying a hashtag; the quotes are usually necessary
    250 "%212998617"         posts tagged with a numeric location id
    251 "@profile_name"      the profiles that profile_name follows, i.e. its followees
    252 ```
    253 
    254 Location IDs are the non-obvious pivot: they are numeric, they are stable, and
    255 [instagram-location-search](#instagram-location-search) below turns a coordinate into a list of
    256 them, which turns "somewhere near this junction" into a set of addressable targets.
    257 
    258 ## X / Twitter
    259 
    260 Heavily restricted since the API changes. Advanced search still works while logged in, and is
    261 still the best tool on the platform — the operator set is below under
    262 [Key tools](#key-tools).
    263 
    264 Archived snapshots are now often more reliable than the live site for deleted content — go to the
    265 [Wayback Machine](https://web.archive.org/) first.
    266 
    267 One thing still works with no session at all: a status ID dates itself. If you have a screenshot
    268 with a URL in it, you have the post's creation time whether or not the post survives:
    269 
    270 ```bash
    271 # status id -> UTC creation time, from the Snowflake layout above
    272 python3 - <<'PY'
    273 import datetime
    274 for sid in (1949394974262350098, 1234567890123456789):
    275     ms = (sid >> 22) + 1288834974657
    276     print(sid, datetime.datetime.fromtimestamp(ms / 1000, datetime.timezone.utc).isoformat())
    277 PY
    278 ```
    279 
    280 ```text
    281 1949394974262350098 2025-07-27T09:02:34.528000+00:00
    282 1234567890123456789 2020-03-02T19:54:56.824000+00:00
    283 ```
    284 
    285 That is a creation time for the ID, which for an original post is the post's time. For a quote-post
    286 or a reply it is that reply's time, not the thing it replies to, and the `conversation_id` operator
    287 below is how you get back to the root.
    288 
    289 ## TikTok
    290 
    291 - Profiles and individual videos are viewable without an account.
    292 - [Bellingcat's TikTok Hashtag Analysis](https://github.com/bellingcat/tiktok-hashtag-analysis)
    293   collects posts by hashtag and charts co-occurrence, which is the useful bulk primitive.
    294 - Video download and metadata extraction both work through `yt-dlp`.
    295 
    296 Dating a TikTok needs neither an account nor a live video, because the ID carries the time. This is
    297 the whole of what [Bellingcat's extractor](https://bellingcat.github.io/tiktok-timestamp) does, and
    298 doing it locally means it also works for a URL you only have as text:
    299 
    300 ```bash
    301 # pull the id out of the URL and decode it -- no network request at all
    302 url='https://www.tiktok.com/@user/video/7169068344567909638'
    303 python3 -c "
    304 import datetime, re, sys
    305 vid = int(re.search(r'/video/(\d+)', sys.argv[1]).group(1))
    306 print(vid, datetime.datetime.fromtimestamp(vid >> 32, datetime.timezone.utc).isoformat())
    307 " "$url"
    308 
    309 # cross-check against what the platform says, which needs the video to still exist
    310 yt-dlp --dump-json "$url" | jq '{id,timestamp,upload_date,uploader_id,view_count}'
    311 ```
    312 
    313 ```text
    314 7169068344567909638 2022-11-23T04:46:37+00:00
    315 ```
    316 
    317 When the decoded time and `yt-dlp`'s `timestamp` disagree by more than a second or two, the
    318 interesting case is a reupload: the ID in the URL belongs to the copy you are looking at, and an
    319 earlier copy with an earlier ID is what you are actually hunting for.
    320 
    321 ## YouTube
    322 
    323 Still the most researcher-friendly large platform: metadata, channel listings, comments and
    324 auto-generated subtitles all come out of `yt-dlp`, covered below.
    325 
    326 Upload timestamps are in the metadata and are a reliable earliest-possible date for footage.
    327 
    328 The two invocations that do most of the work on this platform are a channel listing and a subtitle
    329 dump — one tells you what exists, the other turns hours of video into text you can grep:
    330 
    331 ```bash
    332 # the whole channel as a dated listing, without touching a single video file
    333 yt-dlp --flat-playlist --dump-json 'https://youtube.com/@channel/videos' \
    334   | jq -r '[.upload_date, .id, .duration, .title] | @tsv' | sort
    335 
    336 # only uploads inside a window, which is how you tie output to an event
    337 yt-dlp --flat-playlist --dateafter 20260101 --datebefore 20260331 \
    338   -O '%(upload_date)s %(id)s %(title)s' 'https://youtube.com/@channel/videos'
    339 
    340 # subtitles for every video in a playlist, as text, no media
    341 yt-dlp --write-auto-subs --sub-langs 'en.*' --convert-subs srt --skip-download \
    342   -o '%(upload_date)s-%(id)s.%(ext)s' 'https://youtube.com/playlist?list=PLID'
    343 
    344 # then the actual search, across the lot
    345 grep -rin 'phrase you care about' *.srt
    346 ```
    347 
    348 `--flat-playlist` is what keeps a channel listing cheap: it reads the playlist entries without
    349 resolving each video, so a 2,000-video channel is one request chain rather than 2,000. The price is
    350 that some fields are absent from flat entries, so confirm a field exists in `--dump-json` output
    351 before you build a pipeline on it.
    352 
    353 ## Reddit
    354 
    355 - Profiles, posts and comment trees are readable without an account, and appending `.json` to any
    356   Reddit URL returns the same thing as structured data.
    357 - Deleted and removed content is the point of interest, and Reddit does not serve it. The
    358   historical archives do — Arctic Shift below is the live successor to Pushshift.
    359 - Subreddit moderator lists, wiki pages and automoderator configs are public and frequently
    360   forgotten by the people who wrote them.
    361 
    362 The `.json` trick, with the parameters that matter:
    363 
    364 ```bash
    365 # any Reddit URL plus .json, with escaping turned off and a full page
    366 curl -s -A 'research-tooling/1.0 (contact: you@example.test)' \
    367   'https://www.reddit.com/user/some_user/comments/.json?limit=100&raw_json=1' \
    368   | jq -r '.data.children[] | "\(.data.created_utc)  \(.data.subreddit)  \(.data.body[0:100])"'
    369 
    370 # page with the fullname the previous response handed back
    371 curl -s -A 'research-tooling/1.0' \
    372   'https://www.reddit.com/r/subreddit/new/.json?limit=100&raw_json=1&after=t3_abc123'
    373 ```
    374 
    375 `raw_json=1` is not cosmetic: without it Reddit replaces `<`, `>` and `&` with HTML entities in
    376 every string field, so a URL or a quoted snippet comes back mangled and your `grep` misses it.
    377 Expect this route to fail anyway. Reddit now returns **403 to unauthenticated JSON requests from
    378 datacentre address ranges regardless of User-Agent** — tested from a hosted address, a custom agent
    379 string makes no difference — so a VPS or CI runner gets nothing while a residential browser
    380 session gets everything. That asymmetry is exactly why the archive below is not optional.
    381 
    382 ## Bluesky
    383 
    384 Structurally the most open platform here, because the AT Protocol is a public API by design: every
    385 post, follow and profile is a record in a repository you can read without an account.
    386 
    387 - **DIDs are permanent.** A handle is a pointer; the `did:plc:…` identifier survives every rename,
    388   and the PLC directory keeps an audit log of every change to it — including the account's first
    389   operation, which dates its creation.
    390 - **Follow graphs are fully enumerable**, which makes network analysis possible in a way it is not
    391   anywhere else on this page.
    392 - **Search is the one thing that is not open.** Post search now needs an authenticated session;
    393   profile and feed reads do not.
    394 
    395 There are two layers and it is worth knowing which one you are on. The AppView
    396 (`public.api.bsky.app`) serves the rendered, moderated view that the app shows. The account's own
    397 PDS serves the raw repository, which includes record types the app never displays:
    398 
    399 ```bash
    400 # which record collections this account actually holds -- not just posts
    401 curl -s 'https://bsky.social/xrpc/com.atproto.repo.describeRepo?repo=atproto.com' \
    402   | jq '{handle, did, handleIsCorrect, collections}'
    403 
    404 # raw records from one collection, straight out of the repository
    405 curl -s 'https://bsky.social/xrpc/com.atproto.repo.listRecords?repo=atproto.com&collection=app.bsky.feed.post&limit=100&reverse=true' \
    406   | jq -r '.records[] | "\(.value.createdAt)  \(.uri | split("/") | last)  \(.value.text[0:80])"'
    407 ```
    408 
    409 `reverse=true` on `listRecords` gives you the **oldest** records first, which is how you read an
    410 account's first ever posts without paging through everything since. The full tool coverage, the
    411 AppView endpoints and the `goat` CLI are below under [Key tools](#key-tools).
    412 
    413 ## Key tools
    414 
    415 ### yt-dlp
    416 
    417 Does far more than download: it reads the platform's own metadata for a video, a channel or a
    418 playlist, and it covers well over a thousand sites, so the same invocation works on YouTube,
    419 TikTok, Twitch, Instagram and most of the long tail. For dating footage and enumerating a
    420 channel's output it is the first tool to reach for.
    421 
    422 ```bash
    423 pipx install yt-dlp
    424 
    425 # everything the platform will tell you about one video, without downloading it
    426 yt-dlp --dump-json 'https://youtube.com/watch?v=VIDEOID' | jq '{upload_date,uploader_id,duration,view_count}'
    427 
    428 # just the fields you need, printed — faster to read than the full JSON
    429 yt-dlp -O '%(upload_date)s %(uploader_id)s %(title)s' 'https://youtube.com/watch?v=VIDEOID'
    430 
    431 # the whole channel as a listing, without touching the videos themselves
    432 yt-dlp --flat-playlist --dump-json 'https://youtube.com/@channel/videos' \
    433   | jq -r '[.upload_date,.id,.title] | @tsv'
    434 
    435 # the entire playlist as ONE json object, for feeding a script rather than a reader
    436 yt-dlp --flat-playlist --dump-single-json 'https://youtube.com/@channel/videos' \
    437   | jq '{channel: .uploader, count: (.entries | length)}'
    438 
    439 # auto-generated subtitles, which turn hours of video into greppable text
    440 yt-dlp --write-auto-subs --sub-langs en --skip-download 'https://youtube.com/watch?v=VIDEOID'
    441 
    442 # which subtitle tracks exist at all, before you guess a language code
    443 yt-dlp --list-subs 'https://youtube.com/watch?v=VIDEOID'
    444 
    445 # comments, saved alongside the metadata, for identifying participants
    446 yt-dlp --write-info-json --write-comments --skip-download 'https://youtube.com/watch?v=VIDEOID'
    447 
    448 # the thumbnail set, which is a reverse-image-search input and sometimes an earlier frame
    449 yt-dlp --write-thumbnail --skip-download 'https://youtube.com/watch?v=VIDEOID'
    450 
    451 # one clip out of a long stream, by timestamp, instead of the whole six hours
    452 yt-dlp --download-sections '*01:12:30-01:14:00' 'https://youtube.com/watch?v=VIDEOID'
    453 
    454 # a date window across a channel, for tying output to an event
    455 yt-dlp --flat-playlist --dateafter 20260101 --datebefore 20260331 \
    456   -O '%(upload_date)s %(id)s' 'https://youtube.com/@channel/videos'
    457 
    458 # filter on any metadata field rather than on dates; "&" is how conditions AND together
    459 yt-dlp --match-filters 'duration > 600 & view_count > 10000' \
    460   -O '%(id)s %(duration)s %(view_count)s' 'https://youtube.com/@channel/videos'
    461 
    462 # "?" after the operator keeps items whose field is absent, instead of dropping them silently
    463 yt-dlp --match-filters 'view_count >? 10000' \
    464   -O '%(id)s %(view_count)s' 'https://youtube.com/@channel/videos'
    465 
    466 # a TikTok video with its metadata — same tool, same flags
    467 yt-dlp --write-info-json 'https://www.tiktok.com/@user/video/1234567890123456789'
    468 
    469 # items 1 to 50 of a playlist, with an archive file so a resumed run does not re-fetch
    470 yt-dlp -I 1:50 --download-archive seen.txt 'https://youtube.com/playlist?list=PLID'
    471 
    472 # a livestream from its start, for an event already in progress
    473 yt-dlp --live-from-start 'https://youtube.com/watch?v=LIVEID'
    474 
    475 # the logged-in surface, when a public one does not exist — uses your browser's session
    476 yt-dlp --cookies-from-browser firefox --dump-json URL
    477 ```
    478 
    479 `upload_date` is the platform's own value and is a reliable earliest-possible date for footage;
    480 it is not the date the footage was shot. Fields vary by extractor, so check with `--dump-json`
    481 before scripting against a key — `--match-filters` silently matches nothing when the field you
    482 named does not exist on that extractor, which looks identical to a channel with no matching
    483 videos. The `>?` form above is the documented escape: it treats an absent field as a pass, so a
    484 sweep that returns nothing is telling you about the videos rather than about the extractor. Passing `--cookies-from-browser` attaches your real session to every request, which
    485 attributes the activity to that account — use a research profile. Extractors break when platforms
    486 change, so `pipx upgrade yt-dlp` before blaming the URL.
    487 
    488 ### gallery-dl
    489 
    490 The image-and-gallery counterpart to yt-dlp: it walks a profile, a tag or a thread and takes the
    491 media plus the per-item JSON. Where it earns its place is bulk profile capture with metadata
    492 intact — the captions, timestamps and IDs that make a media set evidence rather than a folder of
    493 pictures.
    494 
    495 ```bash
    496 pipx install gallery-dl
    497 
    498 # what it would fetch, and the metadata for each item, without downloading
    499 gallery-dl -j 'https://www.instagram.com/username/'
    500 
    501 # the available metadata keys for a URL, so your filters reference real fields
    502 gallery-dl -K 'https://www.reddit.com/user/username/submitted/'
    503 
    504 # media plus a JSON sidecar per file, into a named directory
    505 gallery-dl --write-metadata -D ./capture/username 'https://twitter.com/username/media'
    506 
    507 # direct URLs only, to hand to an archiving tool instead of downloading yourself
    508 gallery-dl -g 'https://www.reddit.com/r/subreddit/comments/abc123/'
    509 
    510 # a chosen field per item, printed as the run goes, instead of parsing json afterwards
    511 gallery-dl --no-download -N '{date} {num} {filename}' 'https://www.reddit.com/user/username/submitted/'
    512 
    513 # the first 50 items of a long profile, which is usually all you need to characterise it
    514 gallery-dl --range 1-50 'https://www.flickr.com/photos/username/'
    515 
    516 # only the large images, skipping avatars and thumbnails
    517 gallery-dl --filter 'image_width >= 1000' 'https://example-gallery.test/user/x'
    518 
    519 # an archive file, so re-running a monitored profile fetches only what is new
    520 gallery-dl --download-archive seen.sqlite3 'https://www.instagram.com/username/'
    521 
    522 # many targets from a file, which is how a watchlist gets run
    523 gallery-dl -i targets.txt --download-archive seen.sqlite3
    524 
    525 # slow and polite: the alternative is a 429 and then a block
    526 gallery-dl --sleep 3 --sleep-request 2 --sleep-429 120 'https://www.instagram.com/username/'
    527 
    528 # hash each file as it lands, so the capture is checkable later
    529 gallery-dl --exec 'sha256sum {} >> hashes.txt' -D ./capture 'https://twitter.com/username/media'
    530 
    531 # keep the raw HTML the extractor parsed, which is the only way to debug an empty run
    532 gallery-dl --write-pages -j 'https://www.instagram.com/username/'
    533 
    534 # a logged-in session where the public surface is gone
    535 gallery-dl --cookies-from-browser firefox 'https://www.instagram.com/username/'
    536 ```
    537 
    538 Per-site options live in a config file (`gallery-dl --config-create` writes a starter), and most
    539 rate-limit problems are solved there rather than on the command line — `-o KEY=VALUE` sets one
    540 inline when you do not want to edit the file. Timestamps in the sidecar JSON are the platform's,
    541 which is the point — but the field name differs per extractor, so read `-K` output before you build
    542 a timeline. A profile capture is a snapshot: items deleted before you ran it are simply absent, and
    543 nothing in the output tells you that. An extractor that returns zero items is far more often broken
    544 than correct, and `--write-pages` is how you tell the difference.
    545 
    546 ### Instaloader
    547 
    548 Instagram specifically, and still working where most Instagram tooling has died. Public profiles,
    549 hashtags, locations and stories, with geotags and comments where they exist.
    550 
    551 ```bash
    552 pipx install instaloader
    553 
    554 # a public profile: posts, captions and metadata JSON
    555 instaloader profile_name
    556 
    557 # metadata only — no media, much faster, much less conspicuous
    558 instaloader --no-pictures --no-videos profile_name
    559 
    560 # comments and geotags, which are the parts worth having
    561 instaloader --comments --geotags profile_name
    562 
    563 # just the profile picture at full resolution, for reverse image search
    564 instaloader --no-posts profile_name
    565 
    566 # cap a location run; --count does not apply to ordinary profile targets
    567 instaloader --count 50 --no-videos '%212998617'
    568 
    569 # posts within a date window, using a Python filter expression
    570 instaloader --post-filter 'date_utc >= datetime(2026,1,1)' profile_name
    571 
    572 # only posts that carry a location, which is the geolocation-relevant subset
    573 instaloader --post-filter 'location is not None' --geotags profile_name
    574 
    575 # a hashtag rather than an account
    576 instaloader '#hashtag'
    577 
    578 # a numeric location id, which is what instagram-location-search below produces
    579 instaloader '%212998617'
    580 
    581 # everywhere this account has been tagged by other people
    582 instaloader --tagged --no-posts profile_name
    583 
    584 # one post by shortcode; the leading dash needs the -- separator first
    585 instaloader -- -B_Nmp6MlQGV
    586 
    587 # readable JSON instead of xz-compressed, for grepping a capture afterwards
    588 instaloader --no-compress-json --no-pictures --no-videos profile_name
    589 
    590 # stories and highlights need a session; both attribute the view to that account
    591 instaloader --login YOUR_RESEARCH_ACCOUNT --stories --highlights profile_name
    592 
    593 # reuse a browser session instead of handing credentials to the tool
    594 instaloader --load-cookies firefox --sessionfile ./research.session profile_name
    595 
    596 # stop on the first rate-limit response rather than grinding into a block
    597 instaloader --abort-on 429,401 --max-connection-attempts 1 profile_name
    598 
    599 # resume by stopping at the first already-downloaded post
    600 instaloader --fast-update profile_name
    601 
    602 # resume by recorded timestamp instead, which is the only resume a metadata-only run can use
    603 instaloader --no-pictures --no-videos --latest-stamps stamps.ini profile_name
    604 ```
    605 
    606 **Viewing a story is attributed.** The account whose session you used appears in the poster's
    607 viewer list, so `--stories` is never a passive operation. Unauthenticated use is rate-limited
    608 aggressively and then blocked by IP; authenticated use gets the account flagged and sometimes
    609 disabled — `--abort-on 429` is the difference between one bad response and a dead account.
    610 `--no-pictures` cannot be combined with `--fast-update`, because the resume logic keys off
    611 downloaded media, so a metadata-only monitoring run needs `--latest-stamps` instead. Geotags are
    612 only present where the poster added them, and the location name is user-chosen, not GPS.
    613 
    614 ### instagram-location-search
    615 
    616 Bellingcat's answer to "what is Instagram's own name for this place". It takes a coordinate and
    617 returns the location tags Instagram knows about nearby, with their numeric IDs — which is the
    618 pivot, because an ID is an addressable target for Instaloader while a place name is not.
    619 
    620 ```bash
    621 pip install instagram-location-search
    622 
    623 # a coordinate to a list of location tags, as CSV
    624 instagram-location-search --cookie "$IG_COOKIE" --lat 51.5074 --lng -0.1278 --csv locs.csv
    625 
    626 # every output format at once: the raw API response, a geospatial layer, and a map to eyeball
    627 instagram-location-search --cookie "$IG_COOKIE" --lat 51.5074 --lng -0.1278 \
    628   --json locs.json --geojson locs.geojson --map locs.html
    629 
    630 # just the IDs, which is the form the next tool wants
    631 instagram-location-search --cookie "$IG_COOKIE" --lat 51.5074 --lng -0.1278 --ids locs.txt
    632 
    633 # then feed each one to Instaloader as a location target
    634 while read -r id; do instaloader --no-videos --geotags "%$id"; done < locs.txt
    635 ```
    636 
    637 It needs a session cookie, so every search is attributed to that account. The locations it returns
    638 are *Instagram's* places, which are user-created: there are duplicates, misspellings, joke entries
    639 and venues that have closed, and a radius search returns them by proximity to the coordinate rather
    640 than by relevance. A location tag on a post means the poster selected that place from a list, not
    641 that a GPS fix put them there — see [Geolocation & Chronolocation](/sheets/osint/geolocation) for
    642 what actually constitutes placing an image.
    643 
    644 ### Telegram: Telepathy and Telethon
    645 
    646 [Telepathy](https://github.com/proseltd/Telepathy-Community) is the toolkit the Telegram sections
    647 of older OSINT guides point at, but it has been **unmaintained since July 2024** and the current
    648 source branch does not install cleanly. The commands below document the published 2.3.4 interface;
    649 for anything you need to keep working, the durable route is Telethon, which tracks the Telegram
    650 API itself.
    651 
    652 ```bash
    653 pip3 install telepathy
    654 
    655 telepathy -t channelname                  # basic channel scan
    656 telepathy -t channelname -c               # comprehensive: adds the message history archive
    657 telepathy -t channelname -c -f            # plus a forward edgelist, the useful part
    658 telepathy -t channelname -c -m            # media archiving alongside the comprehensive scan
    659 telepathy -t channelname -c -r            # archive channel replies as well as posts
    660 telepathy -t username -u                  # look up a user your account has encountered
    661 telepathy -e                              # export your account's chat list to CSV
    662 telepathy -t channelname -a 2             # run from an alternative number or API key set
    663 ```
    664 
    665 Both need API credentials from [my.telegram.org](https://my.telegram.org/), which are tied to a
    666 real phone number — use one you are willing to lose. Telethon in forty lines does the part that
    667 matters, and does not rot:
    668 
    669 ```python
    670 # pip install telethon
    671 from collections import Counter
    672 from telethon.sync import TelegramClient
    673 from telethon.errors import FloodWaitError
    674 from telethon.tl.functions.channels import GetFullChannelRequest
    675 from telethon.tl.types import InputMessagesFilterPhotos
    676 
    677 with TelegramClient('research-session', API_ID, API_HASH) as client:
    678     entity = client.get_entity('channelname')
    679     print(entity.id, entity.title, entity.date)        # channel id and creation date
    680 
    681     # what Telegram itself reports about the channel, including the linked discussion group
    682     full = client(GetFullChannelRequest(entity)).full_chat
    683     print(full.participants_count, full.linked_chat_id, (full.about or '')[:120])
    684 
    685     # message history, newest first; message IDs are sequential, so gaps are deletions
    686     for msg in client.iter_messages(entity, limit=500):
    687         fwd = msg.forward.chat.username if msg.forward and msg.forward.chat else None
    688         print(msg.id, msg.date.isoformat(), fwd or '-', (msg.message or '')[:80])
    689 
    690     # the count Telegram reports, against the newest id: the difference is deleted content
    691     page = client.get_messages(entity, limit=100)
    692     newest = page[0]
    693     print('reported total', page.total, 'newest id', newest.id, 'gap', newest.id - page.total)
    694 
    695     # oldest first, which is how you read a channel's first posts without paging the lot
    696     for msg in client.iter_messages(entity, limit=20, reverse=True):
    697         print(msg.id, msg.date.isoformat(), (msg.message or '')[:80])
    698 
    699     # server-side keyword search inside one channel -- far cheaper than fetching everything
    700     for msg in client.iter_messages(entity, search='keyword', limit=100):
    701         print(msg.id, msg.date.isoformat(), (msg.message or '')[:120])
    702 
    703     # only messages carrying photos, for building a verification set
    704     for msg in client.iter_messages(entity, filter=InputMessagesFilterPhotos, limit=50):
    705         print(msg.id, msg.date.isoformat(), msg.file.name if msg.file else '-')
    706 
    707     # everything one account posted in a group, which is the behavioural view
    708     for msg in client.iter_messages('groupname', from_user='username', limit=200):
    709         print(msg.id, msg.date.isoformat(), (msg.message or '')[:80])
    710 
    711     # the forward graph: which channels this one amplifies, with counts
    712     sources = Counter(
    713         m.forward.chat.username
    714         for m in client.iter_messages(entity, limit=2000)
    715         if m.forward and m.forward.chat
    716     )
    717     print(sources.most_common(20))
    718 
    719     # membership, where the group exposes it — you appear in their list too
    720     for user in client.iter_participants(entity, limit=200):
    721         print(user.id, user.username, user.first_name)
    722 ```
    723 
    724 Sequential message IDs are the quiet win here: a gap between IDs is deleted content, and the first
    725 ID dates the channel. The `search=` parameter runs server-side, which matters for a channel with
    726 100,000 messages — it is the difference between one request and a thousand. Joining a group to read
    727 it adds your research account to a list other members can see, and scraping at speed from one
    728 account is exactly the behaviour Telegram bans for. Rate limits surface as `FloodWaitError` with a
    729 `seconds` attribute — sleep for exactly that, rather than retrying, because retrying extends it.
    730 
    731 ### telegram-phone-number-checker
    732 
    733 Bellingcat's tool for the one Telegram question with a clean answer: is this phone number
    734 registered, and if so what does the account show. Also takes usernames.
    735 
    736 ```bash
    737 pip install telegram-phone-number-checker
    738 
    739 # credentials go in a .env: API_ID, API_HASH, PHONE_NUMBER
    740 telegram-phone-number-checker --phone-numbers +14155550123
    741 
    742 # several at once, comma-separated
    743 telegram-phone-number-checker --phone-numbers +14155550123,+14155550124
    744 
    745 # pull the profile photos too, for reverse image search
    746 telegram-phone-number-checker --phone-numbers +14155550123 --download-profile-photos
    747 
    748 # by username instead of number
    749 telegram-phone-number-checker --usernames johndoe
    750 
    751 # both in one run, to a named output file
    752 telegram-phone-number-checker --phone-numbers +14155550123 --usernames johndoe --output results.json
    753 
    754 # credentials on the command line, overriding the .env, for a one-off run on a second account
    755 telegram-phone-number-checker --api-id 123456 --api-hash abcdef0123456789 \
    756   --api-phone-number +14155550000 --phone-numbers +14155550123
    757 
    758 # through a proxy, since the lookups come from your account
    759 pip install 'telegram-phone-number-checker[proxy]'
    760 telegram-phone-number-checker --phone-numbers +14155550123 --proxy socks5://127.0.0.1:1080
    761 ```
    762 
    763 Checking a number works by adding it to your account's contacts, which is a write operation on
    764 Telegram's side: the README's own advice is not to use a personal account, and that is not
    765 boilerplate. A negative result means the number is not registered *or* the account's privacy
    766 settings hide it from non-contacts — the tool cannot tell you which. See
    767 [email and phone OSINT](/sheets/osint/email-and-phone) for the rest of the number work.
    768 
    769 ### tiktok-hashtag-analysis
    770 
    771 Bellingcat's bulk primitive for TikTok: collect the posts carrying one or more hashtags, store the
    772 metadata, and chart which hashtags co-occur. That co-occurrence graph is how a coordinated
    773 campaign becomes visible, which individual video analysis never shows.
    774 
    775 ```bash
    776 pip install tiktok-hashtag-analysis
    777 
    778 # collect posts for one or more hashtags
    779 tiktok-hashtag-analysis london paris
    780 
    781 # hashtags from a file, when the list is long
    782 tiktok-hashtag-analysis --file hashtags.txt --output-dir ./tiktok-run
    783 
    784 # the top 20 co-occurring hashtags, as a table
    785 tiktok-hashtag-analysis london --number 20 --table
    786 
    787 # the same, plotted
    788 tiktok-hashtag-analysis london --number 20 --plot
    789 
    790 # download the videos too, capped per hashtag
    791 tiktok-hashtag-analysis london --download --limit 200
    792 
    793 # short forms, for the same thing: -d download, -t table, -p plot
    794 tiktok-hashtag-analysis london -d -t --number 20 --limit 50
    795 
    796 # keep a log file, which is the only record of what a long run actually did
    797 tiktok-hashtag-analysis london --log ./run.log --output-dir ./tiktok-run
    798 
    799 # watch the browser work, which is how you debug a run that returns nothing
    800 tiktok-hashtag-analysis london --headed -v
    801 ```
    802 
    803 It drives a real browser, so it is slow and it breaks when TikTok changes its front end — a run
    804 that returns zero posts for a busy hashtag means the scraper is broken, not that the hashtag is
    805 empty. Check with `--headed` before concluding anything. Hashtag collection is a sample, never the
    806 complete set, so counts are comparative at best, and two runs an hour apart will not agree. For a
    807 single video's exact upload time, decode the ID as shown under [TikTok](#tiktok) above, and
    808 `yt-dlp --dump-json` gets the rest.
    809 
    810 ### Arctic Shift
    811 
    812 The live successor to Pushshift for Reddit: a public, keyless archive of posts and comments,
    813 including content since deleted from Reddit itself. This is the only practical way to read a
    814 removed comment or reconstruct a user's history after they wiped it.
    815 
    816 ```bash
    817 AS='https://arctic-shift.photon-reddit.com/api'
    818 
    819 # everything an author posted, oldest first
    820 curl -s -G "$AS/posts/search" --data-urlencode 'author=some_user' \
    821   --data-urlencode 'sort=asc' --data-urlencode 'limit=100' \
    822   | jq -r '.data[] | "\(.created_utc)  \(.subreddit)  \(.title)"'
    823 
    824 # their comments, which are usually more revealing than their posts
    825 curl -s -G "$AS/comments/search" --data-urlencode 'author=some_user' \
    826   --data-urlencode 'limit=100' | jq -r '.data[] | "\(.subreddit)  \(.body[0:120])"'
    827 
    828 # keyword search inside a subreddit and a date window
    829 curl -s -G "$AS/posts/search" --data-urlencode 'subreddit=worldnews' \
    830   --data-urlencode 'title=wuhan' --data-urlencode 'after=2019-12-30' --data-urlencode 'limit=10'
    831 
    832 # title and selftext together, which is what 'query' is for
    833 curl -s -G "$AS/posts/search" --data-urlencode 'subreddit=osint' \
    834   --data-urlencode 'query=telegram' --data-urlencode 'limit=25'
    835 
    836 # only the fields you want, which keeps a wide sweep manageable
    837 curl -s -G "$AS/posts/search" --data-urlencode 'subreddit=osint' \
    838   --data-urlencode 'fields=title,created_utc,author' --data-urlencode 'limit=100'
    839 
    840 # posts linking a specific domain — how a site spreads across subreddits
    841 curl -s -G "$AS/posts/search" --data-urlencode 'url=example-news-daily.com' \
    842   --data-urlencode 'limit=100' | jq -r '.data[].subreddit' | sort | uniq -c | sort -rn
    843 
    844 # posting cadence rather than content: a month-by-month count
    845 curl -s -G "$AS/posts/search/aggregate" --data-urlencode 'subreddit=osint' \
    846   --data-urlencode 'aggregate=created_utc' --data-urlencode 'frequency=month' \
    847   --data-urlencode 'after=2026-01-01' | jq -r '.data[] | "\(.created_utc[0:7])  \(.count)"'
    848 
    849 # where an account actually operates, weighted and ranked
    850 curl -s -G "$AS/users/interactions/subreddits" --data-urlencode 'author=some_user' \
    851   --data-urlencode 'limit=10' | jq -r '.data[] | "\(.count)\t\(.subreddit)"'
    852 
    853 # who an account talks to, which is the social graph Reddit does not expose
    854 curl -s -G "$AS/users/interactions/users" --data-urlencode 'author=some_user' \
    855   --data-urlencode 'min_count=3' --data-urlencode 'limit=20'
    856 
    857 # every wiki page a subreddit has, including the ones nobody links to
    858 curl -s -G "$AS/subreddits/wikis/list" --data-urlencode 'subreddit=osint' | jq -r '.data[]'
    859 
    860 # the full comment tree under one post, up to 25,000 comments
    861 curl -s -G "$AS/comments/tree" --data-urlencode 'link_id=abc123' \
    862   --data-urlencode 'limit=9999'
    863 
    864 # known IDs straight to records, up to 500 per call
    865 curl -s -G "$AS/comments/ids" --data-urlencode 'ids=t1_aaa,t1_bbb'
    866 
    867 # accounts by username prefix, for a sockpuppet family
    868 curl -s -G "$AS/users/search" --data-urlencode 'author_prefix=civicwatch' \
    869   --data-urlencode 'sort_type=author' --data-urlencode 'limit=100' | jq -r '.data[].author'
    870 
    871 # an RSS feed of a standing query, for monitoring rather than a one-off sweep
    872 curl -s -G "$AS/comments/search" --data-urlencode 'subreddit=osint' \
    873   --data-urlencode 'format=rss' --data-urlencode 'limit=25'
    874 ```
    875 
    876 A real response, from the aggregate query above:
    877 
    878 ```json
    879 {"data":[{"created_utc":"2025-12-31T23:00:00.000Z","count":"390"},
    880          {"created_utc":"2026-01-31T23:00:00.000Z","count":"376"},
    881          {"created_utc":"2026-02-28T23:00:00.000Z","count":"503"},
    882          {"created_utc":"2026-03-31T22:00:00.000Z","count":"450"}]}
    883 ```
    884 
    885 Note the bucket boundaries: `23:00:00Z` then `22:00:00Z` after March, because the buckets are cut
    886 in a local timezone that observes DST, so a month's `created_utc` label is the start of the bucket
    887 and not the month. Keyword search is constrained by design — `title` works only alongside `author`
    888 or `subreddit`, `body` only alongside `author`, `subreddit`, `link_id` or `parent_id` — so a bare
    889 keyword sweep is refused, and it is refused for very active authors and subreddits too. The archive
    890 holds what it ingested at the time, so an edit after ingestion is invisible and a post removed
    891 within seconds of being made may never have been captured. A record here is evidence that text was
    892 published, not that it is still live, and the live check is a separate step.
    893 
    894 ### Bluesky and the AT Protocol
    895 
    896 No key, no account, and a genuinely public API — the most researcher-friendly surface of any
    897 current platform. Everything below runs against the public AppView except post search, which now
    898 requires an authenticated session.
    899 
    900 ```bash
    901 API='https://public.api.bsky.app/xrpc'
    902 
    903 # handle to DID: the identifier that survives renames
    904 curl -s "$API/com.atproto.identity.resolveHandle?handle=bellingcat.com" | jq -r .did
    905 
    906 # the PLC audit log: every handle and server change, and the creation date in entry one
    907 curl -s 'https://plc.directory/did:plc:z72i7hdynmk6r22z27h6tvur/log/audit' \
    908   | jq -r '.[] | "\(.createdAt)  \(.operation.alsoKnownAs // [] | join(","))"'
    909 
    910 # profile: display name, counts, and the account's own description
    911 curl -s "$API/app.bsky.actor.getProfile?actor=bellingcat.com" \
    912   | jq '{handle,displayName,followersCount,postsCount,createdAt}'
    913 
    914 # up to 25 accounts in one request, which is how you compare a suspected cluster
    915 curl -s -G "$API/app.bsky.actor.getProfiles" \
    916   --data-urlencode 'actors=bellingcat.com' --data-urlencode 'actors=atproto.com' \
    917   | jq -r '.profiles[] | "\(.createdAt)  \(.followersCount)\t\(.handle)"'
    918 
    919 # the author's posts, paged with the cursor the response returns
    920 curl -s "$API/app.bsky.feed.getAuthorFeed?actor=bellingcat.com&limit=100" \
    921   | jq -r '.feed[] | "\(.post.indexedAt)  \(.post.record.text[0:100])"'
    922 
    923 # only posts carrying media, which is the usual filter for verification work
    924 curl -s "$API/app.bsky.feed.getAuthorFeed?actor=bellingcat.com&limit=50&filter=posts_with_media"
    925 
    926 # the account's own threads without the replies noise
    927 curl -s "$API/app.bsky.feed.getAuthorFeed?actor=bellingcat.com&limit=50&filter=posts_no_replies"
    928 
    929 # the follow graph, fully enumerable in both directions
    930 curl -s "$API/app.bsky.graph.getFollows?actor=bellingcat.com&limit=100" | jq -r '.follows[].handle'
    931 curl -s "$API/app.bsky.graph.getFollowers?actor=bellingcat.com&limit=100" | jq -r '.followers[].handle'
    932 
    933 # the lists an account curates, which say who it considers relevant
    934 curl -s "$API/app.bsky.graph.getLists?actor=bellingcat.com&limit=50" \
    935   | jq -r '.lists[] | "\(.purpose)\t\(.listItemCount)\t\(.name)"'
    936 
    937 # who amplified one post, and who liked it -- both need the post's AT-URI
    938 curl -s "$API/app.bsky.feed.getRepostedBy?uri=at://did:plc:xxxx/app.bsky.feed.post/yyyy&limit=100" \
    939   | jq -r '.repostedBy[].handle'
    940 curl -s "$API/app.bsky.feed.getLikes?uri=at://did:plc:xxxx/app.bsky.feed.post/yyyy&limit=100" \
    941   | jq -r '.likes[] | "\(.createdAt)  \(.actor.handle)"'
    942 
    943 # a whole thread, replies included, from one post's AT URI
    944 curl -s "$API/app.bsky.feed.getPostThread?uri=at://did:plc:xxxx/app.bsky.feed.post/yyyy"
    945 
    946 # post search needs a session — this returns an auth error on the public host
    947 curl -s -G "$API/app.bsky.feed.searchPosts" --data-urlencode 'q=osint' \
    948   --data-urlencode 'sort=latest' --data-urlencode 'since=2026-01-01' --data-urlencode 'limit=25'
    949 ```
    950 
    951 Once you have a session, `searchPosts` is the richest query surface on the platform: `q` plus
    952 `author`, `mentions`, `domain`, `url`, `lang`, repeated `tag` parameters that AND together, and
    953 `since`/`until` which take a bare `YYYY-MM-DD`. Authenticate with
    954 `com.atproto.server.createSession` against the account's own PDS and send the returned
    955 `accessJwt` as a bearer token:
    956 
    957 ```bash
    958 # a session on the account's own PDS; use an app password, never the real one
    959 PDS='https://bsky.social'
    960 TOKEN=$(curl -s -X POST "$PDS/xrpc/com.atproto.server.createSession" \
    961   -H 'Content-Type: application/json' \
    962   -d '{"identifier":"you.bsky.social","password":"xxxx-xxxx-xxxx-xxxx"}' | jq -r .accessJwt)
    963 
    964 # every post from one account mentioning a domain, inside a date window
    965 curl -s -G "$PDS/xrpc/app.bsky.feed.searchPosts" -H "Authorization: Bearer $TOKEN" \
    966   --data-urlencode 'q=*' --data-urlencode 'author=bellingcat.com' \
    967   --data-urlencode 'domain=example-news-daily.com' --data-urlencode 'since=2026-01-01' \
    968   | jq -r '.posts[] | "\(.record.createdAt)  \(.author.handle)  \(.record.text[0:100])"'
    969 
    970 # posts carrying two tags at once, which is the coordination signal
    971 curl -s -G "$PDS/xrpc/app.bsky.feed.searchPosts" -H "Authorization: Bearer $TOKEN" \
    972   --data-urlencode 'q=*' --data-urlencode 'tag=osint' --data-urlencode 'tag=verification' \
    973   --data-urlencode 'sort=latest' | jq -r '.posts[].author.handle' | sort | uniq -c | sort -rn
    974 ```
    975 
    976 The PLC audit log is the find here: it is an append-only record of every handle the account has
    977 used and every server it has moved between, dated, with the first entry establishing when the
    978 account was created. `indexedAt` is when the relay saw the post, not when the author wrote it —
    979 `record.createdAt` is the author's claim and is trivially spoofable, as is the TID in the record
    980 key, so quote both and say which is which. `searchPosts` is explicitly documented as possibly
    981 non-public, so treat the public host returning results for it as luck rather than a contract, and
    982 note that `since`/`until` filter on the server's `sortAt`, which may not match `createdAt` at all.
    983 Authenticating turns a keyless read into an attributable one.
    984 
    985 ### goat
    986 
    987 The reference AT Protocol CLI, from Bluesky themselves. It does the things `curl` makes awkward:
    988 exporting a whole repository as a single file, decoding record keys, walking the PLC operation log,
    989 and tailing the firehose. For Bluesky work it replaces about ten of the commands above.
    990 
    991 ```bash
    992 brew install goat        # or: go install github.com/bluesky-social/goat@latest
    993 
    994 # resolve an identity to its DID document, including which PDS holds the data
    995 goat resolve wyden.senate.gov
    996 
    997 # every record collection the account holds -- including ones the app never renders
    998 goat ls -c dril.bsky.social
    999 
   1000 # one record, as the author wrote it
   1001 goat get at://dril.bsky.social/app.bsky.feed.post/3kkreaz3amd27
   1002 
   1003 # the entire repository as a CAR file: every record, in one evidential artefact
   1004 goat repo export jay.bsky.team
   1005 
   1006 # and the blobs, which is every image the account has ever uploaded
   1007 goat blob export jay.bsky.team
   1008 
   1009 # the full identity history: handle changes, PDS moves, key rotations, dated
   1010 goat plc history atproto.com
   1011 
   1012 # a record key back to the time the client claims it was written
   1013 goat syntax tid inspect 3kzifvcppte22
   1014 
   1015 # live posts matching nothing in particular, for watching an event unfold
   1016 goat firehose --ops -c app.bsky.feed.post
   1017 
   1018 # handle changes across the whole network, which is a renaming-detection feed
   1019 goat firehose --account-events | jq .payload.handle
   1020 
   1021 # any XRPC query against any host, which is the escape hatch
   1022 goat xrpc query https://public.api.bsky.app app.bsky.actor.getProfile actor==atproto.com
   1023 ```
   1024 
   1025 `goat repo export` is the one to reach for when the account might be deleted: a CAR file is the
   1026 whole repository at one moment, signed, and it keeps working after the account does not. Note the
   1027 `==` in the `xrpc query` form — that is how `goat` distinguishes a query parameter from a header,
   1028 and a single `=` there means something else. Reads need no authentication; `goat account login`
   1029 exists for writes and for the handful of authenticated reads, and logging in attributes everything
   1030 after it. A TID decoded by `syntax tid inspect` reads the clock of whatever generated the record
   1031 key, not the network's, and the TID specification notes that known TIDs can be reused — so it
   1032 dates a claim rather than an event.
   1033 
   1034 ### Meta Ad Library API
   1035 
   1036 The one Meta surface that is genuinely open, and the only reliable way to get spend and reach data
   1037 for political and issue advertising. Searchable by keyword or page, with the creative, the dates,
   1038 the platforms and the demographic breakdown.
   1039 
   1040 ```bash
   1041 # access needs a Meta developer app, ID verification and Ad Library API approval first
   1042 TOKEN="$META_TOKEN"
   1043 G='https://graph.facebook.com/v23.0/ads_archive'
   1044 
   1045 # keyword search; ad_reached_countries is mandatory, and so is search_terms or search_page_ids
   1046 curl -s -G "$G" -d "access_token=$TOKEN" \
   1047   -d 'search_terms=climate' -d "ad_reached_countries=['GB']" \
   1048   -d 'fields=page_name,ad_delivery_start_time,ad_snapshot_url' | jq -r '.data[].page_name' | sort -u
   1049 
   1050 # an exact phrase rather than unordered keywords
   1051 curl -s -G "$G" -d "access_token=$TOKEN" \
   1052   -d 'search_terms=net zero' -d 'search_type=KEYWORD_EXACT_PHRASE' \
   1053   -d "ad_reached_countries=['GB']" -d 'fields=page_name,ad_creative_bodies'
   1054 
   1055 # one page's entire ad history, which is the attribution-relevant query
   1056 curl -s -G "$G" -d "access_token=$TOKEN" \
   1057   -d 'search_page_ids=123456789' -d "ad_reached_countries=['GB']" \
   1058   -d 'ad_active_status=ALL' \
   1059   -d 'fields=ad_creative_bodies,ad_delivery_start_time,ad_delivery_stop_time,spend,impressions'
   1060 
   1061 # who paid, which is the field the transparency rules exist for
   1062 curl -s -G "$G" -d "access_token=$TOKEN" \
   1063   -d 'search_terms=election' -d "ad_reached_countries=['GB']" \
   1064   -d 'fields=bylines,beneficiary_payers,page_name,spend,currency,publisher_platforms'
   1065 
   1066 # where it ran and to whom
   1067 curl -s -G "$G" -d "access_token=$TOKEN" \
   1068   -d 'search_terms=election' -d "ad_reached_countries=['GB']" \
   1069   -d 'fields=delivery_by_region,demographic_distribution,languages'
   1070 
   1071 # only the ads that ran as video, on Instagram rather than Facebook
   1072 curl -s -G "$G" -d "access_token=$TOKEN" \
   1073   -d 'search_terms=election' -d "ad_reached_countries=['GB']" \
   1074   -d 'media_type=VIDEO' -d "publisher_platforms=['INSTAGRAM']" \
   1075   -d 'fields=page_name,ad_snapshot_url,publisher_platforms'
   1076 
   1077 # the non-political categories, which most people never query
   1078 curl -s -G "$G" -d "access_token=$TOKEN" \
   1079   -d 'search_terms=apartment' -d "ad_reached_countries=['GB']" \
   1080   -d 'ad_type=HOUSING_ADS' -d 'fields=page_name,ad_creative_bodies,ad_delivery_start_time'
   1081 
   1082 # a date window, for tying a campaign to an event
   1083 curl -s -G "$G" -d "access_token=$TOKEN" \
   1084   -d 'search_terms=referendum' -d "ad_reached_countries=['IE']" \
   1085   -d 'ad_delivery_date_min=2026-01-01' -d 'ad_delivery_date_max=2026-03-31' \
   1086   -d 'fields=page_name,spend,ad_snapshot_url'
   1087 
   1088 # reach, not impressions: the EU and Brazil figures are per-ad totals
   1089 curl -s -G "$G" -d "access_token=$TOKEN" \
   1090   -d 'search_page_ids=123456789' -d "ad_reached_countries=['IE']" \
   1091   -d 'fields=eu_total_reach,age_country_gender_reach_breakdown,estimated_audience_size'
   1092 
   1093 # page the result set with the cursor Graph hands back
   1094 curl -s -G "$G" -d "access_token=$TOKEN" -d 'search_terms=climate' \
   1095   -d "ad_reached_countries=['GB']" -d 'fields=page_name' -d 'limit=100' -d 'after=CURSOR'
   1096 ```
   1097 
   1098 Leaving out both `search_terms` and `search_page_ids` is a parameter error, not an unfiltered
   1099 dump. `search_terms` matches the ad's own language and is capped at 100 characters, spaces are
   1100 treated as AND unless you set `search_type=KEYWORD_EXACT_PHRASE`, and `search_page_ids` takes at
   1101 most ten IDs. `spend` and `impressions` come back as ranges, not numbers — report them as ranges,
   1102 and note that `eu_total_reach` is a single figure while `estimated_audience_size` is another range.
   1103 Several fields and filters — `bylines`, `delivery_by_region`, `demographic_distribution`,
   1104 `estimated_audience_size_min`/`_max` — exist only for political and issue ads, so an empty result
   1105 for a commercial advertiser means the field does not apply rather than that the ad does not exist.
   1106 Coverage is political and issue ads in the countries where Meta is required to disclose them, plus
   1107 the `EMPLOYMENT_ADS`, `FINANCIAL_PRODUCTS_AND_SERVICES_ADS` and `HOUSING_ADS` categories;
   1108 ordinary commercial ads appear in the web Ad Library with far less metadata and are not in the API
   1109 the same way.
   1110 
   1111 ### X / Twitter advanced search
   1112 
   1113 Web only in practice. The API is priced out of reach for research, `snscrape` has been
   1114 non-functional against X since the 2023 access changes, and the public Nitter instances are
   1115 largely dead — so the live tool is X's own advanced search, which still works while logged in.
   1116 
   1117 ```text
   1118 from:username since:2026-01-01 until:2026-06-30     one account, bounded by date
   1119 to:username                                          replies directed at an account
   1120 "exact phrase" -filter:retweets                      the original post rather than its echoes
   1121 filter:images filter:links                            posts carrying media or outbound links
   1122 filter:replies                                        only replies, which is where arguments live
   1123 url:example-news-daily.com                            who linked a domain
   1124 geocode:51.5074,-0.1278,5km                           posts placed near a coordinate
   1125 min_faves:500 min_retweets:100 min_replies:50         the versions that actually travelled
   1126 lang:ru "phrase"                                      one language, for a multilingual campaign
   1127 (term1 OR term2) -term3                               grouping and negation, as you would expect
   1128 list:12345678 "phrase"                                search inside a list you have built
   1129 conversation_id:1234567890123456789                   one thread, including replies
   1130 from:username filter:media since:2026-03-01           one account's images in one month
   1131 ```
   1132 
   1133 Every one of these requires a logged-in session, which attributes the searching to that account —
   1134 use a research account, and assume the queries are logged. Date-bounded `from:` queries are the
   1135 most reliable form; free-text search silently drops older results, so an empty result for an old
   1136 date range is not evidence of absence. X publishes no machine-readable operator list and its help
   1137 pages block automated fetches, so verify any operator against the advanced-search form itself
   1138 before relying on it — the form is the only authority, and it has quietly lost operators before.
   1139 For deleted posts, go to the [Wayback Machine](https://web.archive.org/) and archive.today first:
   1140 for X specifically, archived snapshots are now a more dependable source than the live site. A
   1141 status ID still dates itself with no session at all, as shown under [X / Twitter](#x--twitter)
   1142 above.
   1143 
   1144 ## Platform tooling at a glance
   1145 
   1146 | Tool | Platform | Account needed | What it gets you |
   1147 | --- | --- | --- | --- |
   1148 | [yt-dlp](https://github.com/yt-dlp/yt-dlp) | 1,000+ sites | No, until the platform insists | Video and channel metadata, subtitles, comments, timed clips |
   1149 | [gallery-dl](https://github.com/mikf/gallery-dl) | Images and galleries, most platforms | No, until the platform insists | Bulk profile capture with per-item JSON sidecars |
   1150 | [Instaloader](https://instaloader.github.io/) | Instagram | No for public posts, yes for stories | Posts, captions, comments, geotags, profile pictures |
   1151 | [instagram-location-search](https://github.com/bellingcat/instagram-location-search) | Instagram | Yes, a session cookie | Location IDs near a coordinate, as CSV, JSON, GeoJSON or a map |
   1152 | [Telethon](https://docs.telethon.dev/) | Telegram | Yes, API credentials and a phone number | Message history, forward graph, member lists, creation dates |
   1153 | [Telepathy](https://github.com/proseltd/Telepathy-Community) | Telegram | Yes, API credentials and a phone number | The same as a CLI, but unmaintained since July 2024 |
   1154 | [Telegram Phone Number Checker](https://github.com/bellingcat/telegram-phone-number-checker) | Telegram | Yes, and it writes to your contacts | Whether a number or username is registered |
   1155 | [tiktok-hashtag-analysis](https://github.com/bellingcat/tiktok-hashtag-analysis) | TikTok | No | Posts by hashtag, plus a co-occurrence table or plot |
   1156 | [Arctic Shift](https://arctic-shift.photon-reddit.com/) | Reddit | No | Deleted posts and comments, wiki page lists, activity aggregates |
   1157 | [goat](https://github.com/bluesky-social/goat) | Bluesky / AT Protocol | No for reads | Repository and blob exports, PLC history, TID decoding, firehose |
   1158 | [Meta Ad Library API](https://developers.facebook.com/docs/graph-api/reference/ads_archive/) | Facebook, Instagram | Yes, an approved developer app | Spend and reach ranges, payer bylines, delivery breakdowns |
   1159 | [X/Twitter Advanced Search](https://x.com/search-advanced) | X | Yes, a logged-in session | Date-, geo- and engagement-bounded post search |
   1160 
   1161 Reach for `yt-dlp` and `gallery-dl` first on any platform either covers, because they are the only
   1162 two on this list maintained at the pace platforms break things. Reach for Arctic Shift and `goat`
   1163 when the question is historical, because both hold data the live platform will not serve. Everything
   1164 else here is worth a check against its last commit date before you trust its output.
   1165 
   1166 ## Tool reference
   1167 
   1168 | Tool | What it does | Cost |
   1169 | --- | --- | --- |
   1170 | [4plebs](https://4plebs.org/) | Searchable archive of specific 4chan boards. Makes it possible to read threads after they are purged from 4chan. | free |
   1171 | [Bellingcat TikTok Date Extract](https://bellingcat.github.io/tiktok-timestamp) | Get the exact upload date + time for tiktok video urls | free |
   1172 | [Bellingcat TikTok Hashtag Analysis](https://github.com/bellingcat/tiktok-hashtag-analysis) | Archive content and metadata from TikTok posts that contain one or more specified hashtags | free |
   1173 | [Blackbird](https://github.com/p1ngul1n0/blackbird) | Check usernames and email addresses on websites and social networks | free |
   1174 | [BskyFollowFinder](https://bsky-follow-finder.theo.io/) | A tool that identifies which Bluesky accounts are followed by a profile’s contacts but not by that profile. Can be used for expanding networks and social… | free |
   1175 | [Disboard](https://disboard.org/servers) | Search for public discord servers | free |
   1176 | [Discord Chat Exporter](https://github.com/Tyrrrz/DiscordChatExporter) | A tool for exporting and backing up Discord chat logs in multiple formats. | free |
   1177 | [DiscordLeaks](https://discordleaks.unicornriot.ninja/) | Search hundreds of thousands of messages leaked from 290+ white-supremacist / nazi discord servers. | free |
   1178 | [F5Bot](https://f5bot.com/) | Sends you an email when a keyword is mentioned on Reddit. | free |
   1179 | [Facebook Video Downloader](http://fdown.net/) | Handy website to download public Facebook videos. Copy paste the URL of the video and download it in the available definition formats. | free |
   1180 | [FindClone](https://findclone.ru/) | Searches images from VK profiles (within certain limits) | free |
   1181 | [Ghunt](https://github.com/mxrch/GHunt) | A command line tool for obtaining information about Google accounts. | free |
   1182 | [Google Account Finder (EPIEOS)](https://tools.epieos.com/google-account.php) | Find the profile picture and public Google Map Reviews + Photos associated with a G-mail adress. Also checks for phone numbers, and checks for email… | free |
   1183 | [Gravatar Email Checker](https://en.gravatar.com/site/check/) | Check if an email address has been used to comment on blogs and whether there is a profile image attached. | free |
   1184 | [HaveIBeenZuckered](https://haveibeenzuckered.com/) | Check if a telephone number is present within the Facebook data breach. | free |
   1185 | [Hoaxy](https://x.com/OSoMe_IU/status/1949394974262350098) | Hoaxy is a web-based search and visualization tool. It helps visualize the spread of information on Bluesky and X (Twitter). | partly free |
   1186 | [Instagram Location Search](https://github.com/bellingcat/instagram-location-search/tree/main) | A command line tool that allows users to find location tags near a specified latitude and longitude. | free |
   1187 | [InstaLoader](https://instaloader.github.io) | Download pictures or videos (with metadata) from Instagram. | free |
   1188 | [Intelligence X Telegram SearchDDD](https://intelx.io/tools?tab=telegram) | Google-based search engine for Telegram (includes Telegago) | free |
   1189 | [LinkdTime](https://github.com/Lucksi/LinkdTime) | Build a clean timeline of any LinkedIn activity from a single URL or a whole list of links. | free |
   1190 | [Meta Content Library](https://transparency.meta.com/researchtools/meta-content-library) | Meta Content Library is a controlled-access tool that lets approved academic and non-profit researchers search the full public archive of Facebook… | free |
   1191 | [MW Geofind](https://mattw.io/youtube-geofind/location) | MW Geofind is a tool designed to help users identify the filming location of YouTube videos, facilitating the exploration of global content from a… | free |
   1192 | [Open Measures](https://public.openmeasures.io) | Open Measures helps open source researchers investigate harmful online activity such as extremism and disinformation. | partly free |
   1193 | [Photo-Map.RU](http://photo-map.ru/) | Geotagged VK posts. | free |
   1194 | [PSNprofiles](https://psnprofiles.com/) | Search PlayStation username, see daily activity, games played, country, and profile pic | free |
   1195 | [RadiTube](https://tool.raditube.com/) | A search engine that searches the subtitles of about 380 (right/left) radical YouTube channels. You query for example for "q says" of "voter fraud" in… | free |
   1196 | [RedditMetis](https://redditmetis.com) | RedditMetis is an online tool that analyses any public Reddit profile and returns charts and statistics on the user’s activity, language, sentiment and… | free |
   1197 | [Sherlock](https://github.com/sherlock-project/sherlock) | Allows a user to search for the presence of specific usernames across more than 400 websites and social networks. | free |
   1198 | [Snap Map](https://map.snapchat.com) | Searchable map of geotagged snaps. | free |
   1199 | [Social Searcher](https://www.social-searcher.com/) | Search hashtags and usernames across various platforms. | partly free |
   1200 | [SteamId.uk](http://steamid.uk/) | Lookup player names, view (more) previously used names, and when accounts befriended eachother (Free). View screenshots of account, (bulk) seach based on… | partly free |
   1201 | [Story Saver](https://storysaver.net) | Download public Instagram Stories, Highlights and Videos. | free |
   1202 | [Strava](https://www.strava.com) | A fitness tracking platform where publicly shared GPS activity data can reveal movement patterns, routines, and precise locations of individuals… | partly free |
   1203 | [Telegago](https://cse.google.com/cse?cx=006368593537057042503:efxu7xprihg) | Telegago is a Google Custom Search Engine tailored for searching public Telegram content for OSINT purposes. | free |
   1204 | [Telegram Group Joiner](https://bellingcat.github.io/telegram-group-joiner/) | Automate joining multiple Telegram groups and channels, ideal for researchers monitoring specific topics. | free |
   1205 | [Telegram Phone Number Checker](https://github.com/bellingcat/telegram-phone-number-checker) | Command line tool for checking if phone numbers are connected to Telegram accounts and retrieving related information where available. | free |
   1206 | [TelegramDB](https://www.telegramdb.org/) | TelegramDB is a searchable database service that allows users to explore public Telegram groups and channels via a dedicated bot. | partly free |
   1207 | [Telemetrio](https://telemetr.io/) | Telemetr.io offers a range of Telegram-related services based on a catalog of Telegram channels: country and category-specific rankings, curated… | partly free |
   1208 | [Telemetry](https://www.telemetryapp.io/) | An analytical search tool for Telegram groups and channels. | partly free |
   1209 | [Telepathy](https://github.com/prose-intelligence-ltd/Telepathy-Community) | Telepathy is a versatile Telegram toolkit for OSINT analysts, enabling chat archiving, memberlist gathering, user location lookup, top poster analysis… | free |
   1210 | [TGStat](https://tgstat.com/) | TGStat is a web-based analytics tool for Telegram that monitors active channels and provides profile analytics and statistics. It tracks channel… | partly free |
   1211 | [TikTok Ad Library](https://library.tiktok.com/ads) | Search and review TikTok ads and related metadata for transparency and research. | free |
   1212 | [TikTok-Api](https://github.com/davidteather/TikTok-Api/) | TikTok-Api is an open-source Python library (unofficial TikTok API) that allows developers and researchers to retrieve various public data from TikTok | free |
   1213 | [tlgrm.eu channels](http://tlgrm.eu/channels) | Search Telegram channels. | free |
   1214 | [Twitter Video Downloader](https://twittervideodownloader.com/) | Download videos from X (formerly Twitter) by converting tweet URLs into downloadable video links. | free |
   1215 | [Twitter/X Location Search](https://twitter.com/explore) | Search for geocoded tweets by their distance from some coordinates. | free |
   1216 | [Vk.watch](http://vk.watch/) | See public comments left by an account, profile photos used, and very basic facial recognition | free |
   1217 | [Who posted what?](https://whopostedwhat.com/) | A tool that allows a keyword search on Facebook on a specific date or within a specific time frame. | free |
   1218 | [X/Twitter Advanced Search](https://x.com/search-advanced) | Twitter/X Advanced Search is X's own tool to help users find more precise information on the platform by filtering posts according to criteria such as… | free |
   1219 | [XboxGamertag](https://xboxgamertag.com/) | Search gamertags, see games played and recorded game clips | free |
   1220 | [YouTube Metadata](https://mattw.io/youtube-metadata/) | An alternative to Amnesty's YT viewer, with slightly more information. | free |
   1221 
   1222 ## Pitfalls
   1223 
   1224 - **Viewing can notify.** LinkedIn tells people who looked. Story views are attributed. Joining a
   1225   Telegram group adds you to a visible list.
   1226 - **Tooling rots fast.** Anything depending on an undocumented endpoint is one deploy from broken.
   1227   Check a tool's commit history before trusting its output.
   1228 - **Deleted is not gone, and present is not original.** Archive first, then analyse.
   1229 - **Reposts dominate.** The account with the most engagement is rarely the origin. Trace back.
   1230 - **An ID dates the record, not the footage.** Decoding a TikTok or Snowflake ID tells you when
   1231   that upload was created. A 2019 clip uploaded in 2026 has a 2026 ID and the arithmetic is still
   1232   right.
   1233 - **A client-written timestamp is a claim.** An atproto TID and `record.createdAt` are both chosen
   1234   by whatever wrote the record; `indexedAt` is the relay's. Quote the one the account did not
   1235   control, and say which you are quoting.
   1236 - **Zero results and broken scraper look identical.** A busy hashtag returning no posts, or a
   1237   `--match-filters` expression naming a field that extractor does not emit, both print nothing.
   1238   Run one query you already know the answer to before you believe a negative.
   1239 - **A pipeline can match and then discard the match.** `grep -o '[0-9]*$'` against
   1240   `data-post="channel/441"` anchors digits to a line that ends in a quote, so it emits a blank
   1241   line per post and exits 0. The extractor worked, the shell threw the answer away, and the
   1242   symptom is an empty file. Count the lines at each stage of a new pipeline before trusting the
   1243   last one.
   1244 - **A date filter may not filter on the date you mean.** Bluesky's `since`/`until` on
   1245   `searchPosts` are documented against the server's `sortAt`, which need not equal
   1246   `record.createdAt`; Arctic Shift's `created_utc` aggregate buckets are cut in a DST-observing
   1247   local zone. Both return a defensible answer to a question you did not ask.
   1248 - **Unauthenticated does not mean unattributed.** Your IP is an identifier, and a sweep from one
   1249   address is a signature. Conversely, authenticating to get past a block converts an anonymous read
   1250   into a logged one tied to an account you will lose.
   1251 - **Datacentre addresses are blocked where browsers are not.** Reddit's `.json` surface returns 403
   1252   from hosted ranges regardless of User-Agent. Code that works on a laptop and fails in CI has not
   1253   found a bug; it has found a policy.
   1254 - **Archives ingest, they do not watch.** Arctic Shift holds what it captured when it captured it:
   1255   an edit made afterwards is invisible, and something deleted within seconds may never have been
   1256   recorded at all.
   1257 - **Spend and reach are ranges.** Meta reports `spend` and `impressions` as buckets. Writing a
   1258   midpoint and calling it a figure invents precision the source does not have.
   1259 - **Counts from a sample are comparative at best.** Hashtag collection, search results and "top
   1260   twenty" lists are samples whose size you do not control. Two runs an hour apart will disagree,
   1261   which is a property of the method, not of the world.
   1262 - **The same handle on four platforms is one lead, not one person.** It takes content, timing or a
   1263   payment record to tie accounts together. Say which of the three you have.
   1264 
   1265 ## Worked example
   1266 
   1267 One datum: the handle **@civicwatchnow**, read off a screenshot of a post, with no platform
   1268 attached and no other context.
   1269 
   1270 1. **Find the platforms.** A username sweep — see
   1271    [username and account discovery](/sheets/osint/usernames-and-accounts) — returns hits on
   1272    Bluesky, YouTube, Reddit and a Telegram channel of the same name. Four surfaces, four different
   1273    amounts of openness.
   1274 2. **Date the accounts, keylessly.** `resolveHandle` gives the Bluesky DID
   1275    `did:plc:z72i7hdynmk6r22z27h6tvur`; the PLC audit log's first entry dates the account to
   1276    **12 March 2026** and shows one earlier handle. An account presenting itself as an established
   1277    watchdog, three months old, is the first real finding.
   1278 3. **Characterise the output.** `yt-dlp --flat-playlist --dump-json` on the YouTube channel lists
   1279    **41 videos**, all uploaded between 2026-04-02 and 2026-05-14 — six weeks — with `upload_date`
   1280    clustered on weekdays. Volume and cadence, from the platform's own metadata.
   1281 4. **Read what was deleted.** Arctic Shift's `comments/search?author=civicwatchnow&limit=100`
   1282    returns **300 comments across eleven subreddits**, of which fourteen no longer exist on Reddit.
   1283    The removed ones carry the same link, which the live profile does not show. The aggregate
   1284    endpoint gives the shape of it:
   1285 
   1286 ```json
   1287 {"data":[{"created_utc":"2026-03-31T22:00:00.000Z","count":"18"},
   1288          {"created_utc":"2026-04-30T22:00:00.000Z","count":"197"},
   1289          {"created_utc":"2026-05-31T22:00:00.000Z","count":"85"}]}
   1290 ```
   1291 
   1292 5. **Map the amplification.** The Telethon forward counter over the Telegram channel's last 2,000
   1293    messages returns a top-twenty list in which **three channels account for 1,340 of 1,612
   1294    forwards**, and the channel's own `entity.date` is **2026-03-14**, two days after the Bluesky
   1295    DID's first PLC operation.
   1296 6. **Check the gaps.** `page.total` reports 1,900 messages against a newest ID of **2,146**, so
   1297    roughly 246 messages — 11 per cent — have been deleted. Which ones is a separate question; that
   1298    they were is not.
   1299 7. **Check for paid distribution.** The Meta Ad Library `search_terms=civicwatch` with
   1300    `ad_reached_countries=['GB']` and `ad_active_status=ALL` returns two advertisers running the
   1301    same creative, with `bylines` naming a company and `spend` as the range **£5,000–£9,999**. That
   1302    is a funding lead the organic surfaces never gave.
   1303 8. **Trace to origin, then archive.** The earliest instance of the clip is the TikTok upload, dated
   1304    from the video ID — `7169068344567909638 >> 32` = **2022-11-23 04:46:37Z** — rather than the
   1305    display text, which read "2 days ago". The YouTube and Telegram copies are later. Capture all
   1306    four profiles, `goat repo export` the Bluesky repository and save the ad snapshots before
   1307    writing anything — see [archiving and evidence](/sheets/osint/archiving-and-evidence).
   1308 
   1309 What you can assert: four accounts under one handle, created within three days of each other in
   1310 March 2026, pushing a single clip whose earliest known copy is dated on TikTok to November 2022,
   1311 amplified by three Telegram channels carrying 83 per cent of the forwards, and promoted by two
   1312 named advertisers with a disclosed spend range. Every date in that sentence came from a
   1313 platform-side identifier or an archive, not from a rendered page.
   1314 
   1315 What would falsify it: an earlier copy of the clip on a platform nobody swept, which would make the
   1316 TikTok ID a reupload's ID rather than the original's; a Telegram channel that was renamed into the
   1317 handle rather than created under it, which the creation date alone cannot distinguish; or an
   1318 advertiser byline that turns out to be a media-buying agency acting for an unrelated client. The
   1319 measurement carrying the most risk is the Bluesky creation date, because a PLC log's first
   1320 operation dates the *DID*, not the project — an operator who registered the DID months before
   1321 using it would push that date earlier with no deception at all, so quote it as "first PLC operation
   1322 on 12 March 2026" rather than as a founding date. What you cannot assert at all is that one person
   1323 runs all four: a shared handle is a lead until content, timing or a payment record ties them, and
   1324 here it is the ad byline, not the handle, doing that work.
   1325 
   1326 ## Broader catalogues
   1327 
   1328 - [Social Media OSINT](https://tools.osintnewsletter.com/tool-categories/social-media-osint)
   1329 
   1330 
   1331 ## More tools
   1332 
   1333 Further tools for this area from the OSINT Newsletter Tools Library ([Facebook](https://tools.osintnewsletter.com/tool-categories/social-media-osint/facebook), [Instagram](https://tools.osintnewsletter.com/tool-categories/social-media-osint/instagram), [Twitter/X](https://tools.osintnewsletter.com/tool-categories/social-media-osint/twitter-x), [Telegram](https://tools.osintnewsletter.com/tool-categories/social-media-osint/telegram), [YouTube](https://tools.osintnewsletter.com/tool-categories/social-media-osint/youtube), [Reddit](https://tools.osintnewsletter.com/tool-categories/social-media-osint/reddit), [Bluesky](https://tools.osintnewsletter.com/tool-categories/social-media-osint/bluesky), [Snapchat](https://tools.osintnewsletter.com/tool-categories/social-media-osint/snapchat), [Other/Multiple Platforms](https://tools.osintnewsletter.com/tool-categories/social-media-osint/other-multiple-platforms)), excluding those already listed above.
   1334 
   1335 | Tool | What it does |
   1336 | --- | --- |
   1337 | [Analyst Research Tools](https://analystresearchtools.com/) | A browser-based investigative research platform pulling together a collection of free lookup tools to help you gather, organise… |
   1338 | [Archivarix Tube Search](https://tube.archivarix.net/) | A free OSINT tool for discovering historical versions of YouTube videos using archived snapshots from archive data. |
   1339 | [Arctic Shift](https://arctic-shift.photon-reddit.com/) | A Reddit search and archival platform that enables investigators to search historical Reddit posts, comments, deleted content… |
   1340 | [Aware Online: Reddit Search](https://www.aware-online.com/en/osint-tools/reddit-search-tool/) | A free OSINT tool that helps investigators search Reddit more effectively by querying posts, comments, usernames, and discussions… |
   1341 | [AzSky](https://azsky.app/) | A Reddit-style interface powered by Bluesky, designed to make exploring decentralised social conversations easier. |
   1342 | [BirdHunt](https://birdhunt.huntintel.io/) | Find tweets based on location, helping you uncover what people were saying in a specific place at a specific time. |
   1343 | [Bluesky Follow Finder](https://bsky-follow-finder.theo.io/) | Surfaces accounts followed by people a profile follows, but that the profile doesn’t. |
   1344 | [BoardReader](https://boardreader.com/) | A specialised search engine for online forums, message boards and discussion communities. |
   1345 | [Buzzglobe](https://buzzglobe.com/) | Social-media search aggregator for discovering publicly indexed profiles, pages, posts and images across multiple platforms. |
   1346 | [de-digger](https://www.dedigger.com/#gsc.tab=0) | A specialist search and discovery engine for finding publicly accessible files on Google Drive. |
   1347 | [Deaddrop](https://deaddrop.theosintconsultants.com/login) | Telegram search engine. |
   1348 | [Discord Leaks](https://discordleaks.unicornriot.ninja/) | A searchable archive of leaked Discord server, RocketChat and Skype messages, mainly from extremist and politically active… |
   1349 | [Dolphin Radar](https://www.dolphinradar.com/) | All-in-one Instagram Tracker & Analyser |
   1350 | [Dumpor](https://dumpor.io/) | Instagram viewer and profile analysis tool that enables anonymous viewing of public Instagram content without logging in. |
   1351 | [Export Comments](https://exportcomments.com/) | Extracts, downloads, and analyses comments from social media platforms and video-sharing sites. |
   1352 | [ExportGram](https://exportgram.net/) | A web tool that lets you extract and download comments from Instagram posts and reels into spreadsheet formats. |
   1353 | [Filmot](https://www.reddit.com/r/languagelearning/comments/odj2gx/ive_built_a_search_engine_across_youtube_captions/) | A search engine that lets you search through YouTube video subtitles and metadata to find videos that contain specific words or… |
   1354 | [Govsky](https://www.govsky.org/) | Allows users to explore and identify official government and public sector accounts on the Bluesky social media platform. |
   1355 | [Graph Tips](https://graph.tips/) | Generates Facebook search queries that recreate much of the functionality of the former Facebook Graph Search. |
   1356 | [Hippie OSINT Toolkit](https://osint.hippie.cat/) | An OSINT web toolkit that allows you to reverse search a domain, a TikTok post, an image, or username (and more). |
   1357 | [Inflact Instagram Profile Viewer](https://inflact.com/instagram-viewer/profile/) | A no-login Instagram viewer that lets you peek at public profiles, posts, and basic engagement data anonymously. |
   1358 | [Instagram Map](https://www.youtube.com/watch?v=R_oOdBTBkGg) | Allows users within Instagram to view geotagged posts, Stories, and shared live locations (where enabled) on a map inside the… |
   1359 | [Internect.info](https://internect.info/) | A search interface for the AT Protocol allowing users to locate and query Bluesky handles and associated account information… |
   1360 | [Nitter](https://nitter.net/) | A privacy-focused alternative front-end for X (Twitter) that lets you view tweets without tracking, ads, or needing an account. |
   1361 | [Sotwe](https://www.sotwe.com/) | Tracks Twitter trends and finds the most popular Twitter users, hashtags and places. |
   1362 | [Sowsearch](https://www.sowsearch.info/) | A specialist Facebook tool to search posts, profiles, comments, places, events, groups, and public content more efficiently using… |
   1363 | [Story Saver (Instagram)](https://www.storysaver.net/en/) | Download public Instagram Stories, highlights and videos. |
   1364 | [Subreddit Stats](https://subredditstats.com/) | A free Reddit analytics tool used to track subreddit growth, activity levels, top posts/users, keyword trends, and community… |
   1365 | [Subreddits.org](https://subreddits.org/) | An online directory and search tool that helps users find and explore Reddit communities (subreddits) by keyword, topic, or… |
   1366 | [Telegago (Telegram)](https://cse.google.com/cse?cx=006368593537057042503:efxu7xprihg#gsc.tab=0) | A custom Google search engine to help users search publicly available content on Telegram channels & groups. |
   1367 | [Telegram Spoiler Decoder](https://spoiler.soxoj.com/) | Instantly recovers hidden spoiler text on MacOS within a Telegram chat. |
   1368 | [ToolsCord](https://toolscord.com/) | A free, browser-based toolkit focused on Discord utilities. |
   1369 | [Treeverse](https://treeverse.app/) | Visualise and explore conversations on Bluesky as interactive threads and networks. |
   1370 | [TweeterID](https://tweeterid.com/) | A simple X/Twitter username and ID conversion tool. |
   1371 | [Twitter Viewer](https://twitterwebviewer.com/) | An online tool that lets you browse Twitter profiles, tweets, and media, and download videos without logging in. |
   1372 | [Waybien](https://waybien.com/en) | A search engine for Telegram, Facebook, Discord, and Whatsapp groups. |
   1373 | [Who Posted What? Facebook Tool](https://whopostedwhat.com/) | A non-public Facebook keyword search tool that allows you to search keywords on specific dates. |
   1374 | [YouTube Video Finder](https://findyoutubevideo.thetechrobo.ca/) | Investigates a YouTube video using its URL or ID, with quick links to archives and analysis tools. |
   1375 
   1376 ## Sources
   1377 
   1378 Both catalogues below are maintained by other people and are considerably larger than
   1379 this page. Use them as the canonical index; this sheet is a working route through them.
   1380 
   1381 - [Bellingcat's Online Investigation Toolkit](https://bellingcat.gitbook.io/toolkit) — ~340 tools, each with its own
   1382   review page covering cost, difficulty, requirements and limitations.
   1383 - [OSINT Newsletter Tools Library](https://tools.osintnewsletter.com) — ~280 tools, organised by investigative goal.
   1384 
   1385 Neither publishes a licence, so nothing here is copied from them: tool names, one-line
   1386 descriptions, cost flags and links are catalogue facts, and the method and commentary are
   1387 this site's own. See [credits](/credits).