social-media-monitoring.md (31354B)
1 --- 2 title: "Cross-Platform Monitoring & Collection" 3 description: "Track accounts, hashtags and narratives across several platforms at once, and watch pages for changes." 4 category: osint 5 subcategory: "Social Media" 6 tags: [osint, monitoring, collection, social-media] 7 tools: [4cat, zeeschuimer, changedetection.io, distill, yt-dlp, rsshub, rss-bridge, arctic-shift] 8 difficulty: intermediate 9 updated: 2026-10-04 10 references: 11 - name: "Bellingcat's Online Investigation Toolkit" 12 url: "https://bellingcat.gitbook.io/toolkit" 13 author: "Bellingcat" 14 license: none 15 relation: derived 16 note: "Tool catalogue: names, descriptions, cost flags and links for this area." 17 - name: "OSINT Newsletter Tools Library" 18 url: "https://tools.osintnewsletter.com" 19 author: "The OSINT Newsletter" 20 license: none 21 relation: derived 22 note: "Second tool catalogue, cross-checked against the above." 23 --- 24 25 ## What this covers 26 27 Watching things over time instead of looking once, and holding what you collect so you can analyse 28 it later without re-collecting. Two distinct jobs with different tooling: **bulk collection** of a 29 corpus for analysis, and **change detection** on a small number of pages you care about. The third 30 thing, which is not optional, is doing either without the pattern of your collection becoming 31 visible to the people you are collecting on. 32 33 ## Method 34 35 1. **Decide what you are watching before you build anything.** A corpus for network analysis and an 36 alert on one bio edit need completely different infrastructure, and building the first when you 37 needed the second is the usual waste. 38 2. **Prefer the platform's own feed.** A native RSS or JSON endpoint is stable, keyless, polite and 39 does not attribute activity to an account. Reach for a scraper only when no feed exists. 40 3. **Store raw, analyse later.** Keep the untouched response alongside anything you derive from it. 41 You will want to re-run the analysis with a different question, and you will not want to 42 re-collect under a worse rate limit. 43 4. **Record gaps as gaps.** A failed poll is a hole in your data, not a quiet day. Log the failures 44 next to the results or your time series will lie to you. 45 5. **Set the cadence to the question.** Hourly is almost never justified. A daily poll that runs 46 for a year is more valuable, and far less conspicuous, than a five-minute poll that gets you 47 blocked in a week. 48 6. **Baseline before you conclude.** Coordination, surges and silences only mean something against 49 what normal looks like for that topic. 50 51 The judgement calls: whether you need the content or only the fact that it changed — the second is 52 far cheaper and far quieter. Whether the collection needs an account at all, because the moment it 53 does, the collection has an identity and a history. And whether you are allowed to keep what you 54 are about to collect, which is a question to settle at the start rather than after you have a 55 database of it. 56 57 ## Collection hygiene 58 59 Monitoring is repeated, scheduled, patterned contact with a target's content. One look is 60 invisible. The same request every fifteen minutes from one address for six months is a signature, 61 and on some platforms it is a signature attached to a logged-in account. 62 63 ```text 64 Account 65 - never your own, and never one that shares a recovery phone or email with your own 66 - a research account per investigation where the platform allows it; one burned 67 account should not cost you the others 68 - assume every logged-in query is retained and attributable, including search terms 69 - some platforms notify a user when a profile is viewed; know which before you look 70 71 Rate 72 - daily is the default. Justify anything faster to yourself in writing 73 - randomise the interval. A poll at exactly :00 every hour is machine-obvious 74 - respect the stated limit, and treat a 429 as a signal to back off for hours, 75 not to retry in a loop 76 - set a real, honest User-Agent that identifies the project. It gets you unblocked 77 more often than a spoofed browser string does 78 79 Network 80 - one address for all of it correlates every collection you run 81 - a residential or VPN exit is a trade: less rate limiting, more attribution risk 82 to whoever pays for it 83 - Tor is blocked by most of these platforms, so it is rarely the answer here 84 85 Footprint 86 - log every request you make, with timestamp and response code. You need it to 87 distinguish "they stopped posting" from "we got blocked" 88 - decide retention up front. Bulk personal data attracts obligations even when 89 every item was public 90 ``` 91 92 The archiving and provenance side of this — hashes, snapshots, what makes a collected item citable 93 later — is on [Archiving & Evidence](/sheets/osint/archiving-and-evidence). Per-platform surfaces, 94 query syntax and the specific scrapers are on 95 [Social Media Platforms](/sheets/osint/social-media-platforms); this sheet is about running those 96 things repeatedly rather than once. 97 98 ## Key tools 99 100 ### 4CAT 101 102 A self-hosted capture-and-analysis platform. It is the serious option because collection and 103 analysis live in the same place: the dataset you capture stays on your machine, and you can re-run 104 a different analysis over it months later without touching the platform again. That is the thing 105 no hosted service gives you. 106 107 ```bash 108 # the documented install is the compose file plus its .env, not a clone 109 mkdir 4cat && cd 4cat 110 curl -O https://raw.githubusercontent.com/digitalmethodsinitiative/4cat/master/docker-compose.yml 111 curl -O https://raw.githubusercontent.com/digitalmethodsinitiative/4cat/master/.env 112 113 # the two settings worth reading before you start it: 114 # SERVER_BIND_ADDRESS=127.0.0.1 localhost only (the shipped default) 115 # PUBLIC_PORT=80 the port the web UI lands on 116 grep -E 'SERVER_BIND_ADDRESS|PUBLIC_PORT|DOCKER_TAG' .env 117 118 docker compose up -d 119 docker compose logs -f backend # first start builds indexes; wait for it 120 121 # the one-time admin link is printed in the logs, not emailed 122 docker compose logs frontend | grep -i 'create a new user\|token' 123 ``` 124 125 Creating a dataset, once it is up: 126 127 ```text 128 1. http://localhost:80 -> "Create dataset". 129 2. Pick the data source. Natively it collects 4chan, 8kun, Bluesky, Telegram, Tumblr 130 and TikTok (from a list of URLs). Everything else arrives as an upload. 131 3. Set the query and the date range. Date range is the field people leave open and 132 then wonder why the job runs for a day -- bound it. 133 4. Submit, then leave it. The backend queues and processes; the dataset appears in 134 your list with a status. Big Telegram or 4chan pulls take hours. 135 5. On the finished dataset, run processors rather than exporting immediately: 136 word frequencies, co-word networks, time series, top posters, image downloads 137 and clustering. Each produces a child dataset you can chain further processors on. 138 6. Export the one you want as CSV or NDJSON, and keep the parent dataset. The parent 139 is the thing you cannot re-create. 140 ``` 141 142 The platform coverage is the live constraint and it changes: X/Twitter, Instagram, LinkedIn, 143 Threads and Pinterest are no longer collected by 4CAT directly — they come in through Zeeschuimer 144 below, or as an upload from elsewhere. Datasets are big, and the Docker volumes will fill a disk 145 without warning, so check free space before a large pull. And a 4CAT capture is a snapshot of what 146 the platform served at that moment: edits and deletions after the capture are invisible to it, 147 which is a feature when you want the original and a trap when you assume it reflects the live site. 148 149 ### Zeeschuimer 150 151 The companion extension, and the current answer to "how do I get X, Instagram or TikTok data into 152 4CAT". It watches the data the platform sends to your browser as you scroll, and keeps it. No API, 153 no scraping requests of its own — it records what your normal browsing already fetched. 154 155 ```text 156 1. Firefox only. Signed .xpi builds are on the project's releases page; Chrome is 157 not supported. 158 2. Install, then open the extension and enable the platform you are about to browse. 159 Supported: TikTok, Instagram, Threads, X/Twitter, Pinterest, Gab, Truth Social, 160 9gag, Imgur, Douyin and RedNote/Xiaohongshu. 161 3. Browse normally -- a profile, a hashtag, a search. The item counter in the 162 extension ticks up as content loads. Content you never scrolled past is not 163 captured, because it was never sent to your browser. 164 4. Export as NDJSON or CSV, or put your 4CAT URL in the extension and upload straight 165 into it as a dataset. 166 5. Record how you browsed: which profile, which tab, how far you scrolled. That is 167 the sampling frame, and without it the dataset's coverage is undocumented. 168 ``` 169 170 This is manual collection with automatic recording, so it does not scale and it does not run 171 unattended — which is also why it survives platform changes better than scrapers do. It captures 172 through your logged-in session, so everything you collect is attributable to that account: use a 173 research profile in a separate Firefox container or profile, never your own. Platform support needs 174 constant maintenance and individual platforms do break; check the releases page before assuming the 175 extension is at fault rather than the site. 176 177 ### changedetection.io 178 179 Self-hosted page watching. This is the one to run when the thing you need to notice is an edit — a 180 bio rewritten, a post deleted, an officer removed from a company page, a policy document quietly 181 amended, a price or a staff list changing. 182 183 ```bash 184 # docker, bound to localhost 185 docker run -d --restart always -p "127.0.0.1:5000:5000" \ 186 -v datastore-volume:/datastore \ 187 --name changedetection.io dgtlmoon/changedetection.io 188 189 # or as a package, if you would rather not run a container 190 pip3 install changedetection.io 191 changedetection.io -d /path/to/empty/data/dir -p 5000 192 ``` 193 194 ```text 195 1. http://localhost:5000 -> "Add new watch", paste the URL. 196 2. Open "Edit" and set the filter before you set anything else. CSS selector, XPath, 197 JSONPath or jq -- point it at the narrowest element carrying the signal. Watching 198 a whole page means alerting on rotating ads, view counters and timestamps. 199 The "Visual Selector" tab picks the element by clicking it. 200 3. Set "Ignore text" for the lines you know churn: relative dates, "N views", 201 cookie-banner text. 202 4. Recheck time: per-watch. Daily for most things. The default applies to every watch 203 you add, so set it low once rather than per watch. 204 5. Notifications: Discord, email, Slack, Telegram or a webhook via Apprise. A webhook 205 into your own notes or ticketing is the one that leaves a record. 206 6. For a page that renders client-side, enable the Playwright/Sockpuppetbrowser 207 fetcher for that watch -- the default plain fetch sees an empty shell and will 208 report "no change" forever. 209 7. "Browser Steps" handles a page behind a form: click, fill, submit, then diff what 210 comes back. 211 8. Every change is stored as a snapshot with a diff view, which is the part that makes 212 it evidence rather than an alert. 213 ``` 214 215 Running it yourself means your watch list is not a third party's business record, which matters 216 when the watch list itself is sensitive. The cost is that a JS-heavy watch runs a real browser and 217 is far heavier than a text diff — a dozen of those on a small VPS will struggle. And a watch only 218 sees what an unauthenticated fetch from your server sees: a page that needs a login, or that 219 geo-varies, needs Browser Steps or will silently watch the wrong thing. 220 221 ### Distill 222 223 The hosted equivalent, for when you will not run infrastructure. Same idea, less setup, and a free 224 tier that is genuinely usable for a handful of watches: **25 monitors total but only 5 in the 225 cloud, a 6-hour minimum cloud interval, 1,000 cloud checks a month, 2 devices, and 30 email alerts 226 a month**. Local monitors in the browser extension are unlimited but only run while the browser is 227 open. 228 229 ```text 230 1. Browser extension or distill.io. On the page you want, click the extension and 231 select the region -- it generates the selector for you. 232 2. Choose local (runs in your browser, unlimited, only while open) or cloud (runs 233 without you, capped as above). For anything that matters, cloud. 234 3. Set the check interval. Anything under 6 hours is a paid feature on the free tier, 235 and six-hourly is adequate for almost all of this work anyway. 236 4. Set the condition, not just "any change" -- Distill supports text conditions, so 237 "alert when the number changes" beats "alert when the page differs". 238 5. Alerts to email or webhook; the email allowance is the binding constraint on free. 239 6. Export the watch list as JSON periodically. It is the only part that is painful to 240 rebuild. 241 ``` 242 243 Your watch list and every snapshot live on their servers, which is the trade for not running 244 anything. For a target that could plausibly subpoena or compromise a third party, that is the wrong 245 trade and `changedetection.io` is the answer instead. The free tier's 6-hour floor also means you 246 will miss a post that goes up and comes down inside a window, which is precisely the kind of 247 deletion worth catching — if that is the scenario, self-host and poll faster. 248 249 ### yt-dlp with a download archive 250 251 For recurring capture of a channel's output, the archive file is the whole trick: it records the ID 252 of everything already fetched, so the next run picks up only what is new. That turns a one-off 253 download into a monitor you can cron. 254 255 ```bash 256 pipx install yt-dlp 257 258 # first run: establish the archive. --break-on-existing stops as soon as it meets 259 # something already recorded, so later runs walk only the new items 260 yt-dlp --download-archive archive.txt --break-on-existing --lazy-playlist \ 261 -o '%(upload_date)s-%(id)s.%(ext)s' \ 262 'https://youtube.com/@channel/videos' 263 264 # metadata-only monitoring: no video files, just the record that it existed 265 yt-dlp --download-archive seen.txt --break-on-existing \ 266 --skip-download --write-info-json --write-thumbnail \ 267 'https://youtube.com/@channel/videos' 268 269 # bound it by date instead, for a backfill of a known window 270 yt-dlp --dateafter 20260101 --datebefore 20260401 --download-archive archive.txt URL 271 272 # subtitles, which turn a channel's output into greppable text as it arrives 273 yt-dlp --download-archive seen.txt --break-on-existing --skip-download \ 274 --write-auto-subs --sub-langs en 'https://youtube.com/@channel/videos' 275 276 # throttle it. these two flags are the difference between a monitor and a nuisance 277 yt-dlp --sleep-requests 2 --sleep-interval 10 --max-sleep-interval 30 \ 278 --download-archive archive.txt URL 279 280 # several channels in one run, each stopping at its own first-seen item 281 yt-dlp --break-per-input --break-on-existing --download-archive archive.txt \ 282 -a channels.txt 283 284 # a cap, so a misconfigured run cannot pull a thousand files overnight 285 yt-dlp --max-downloads 50 --download-archive archive.txt URL 286 ``` 287 288 The archive file is state: back it up, and never delete it to "start fresh" unless you mean to 289 re-download everything. `--break-on-existing` assumes the listing is newest-first, which it is for 290 channels and playlists and is not for some search result pages — on those, drop it and let the 291 archive do the skipping. Extractors break when platforms change their pages, so `pipx upgrade 292 yt-dlp` before blaming a URL, and a cron job that has silently failed for three weeks is worse than 293 no monitoring at all, so alert on non-zero exits. Adding `--cookies-from-browser` attaches a real 294 session to every request and makes the whole monitor attributable; avoid it unless the content 295 genuinely requires a login. Full metadata usage is on 296 [Social Media Platforms](/sheets/osint/social-media-platforms). 297 298 ### Native feeds, before you reach for a bridge 299 300 Several platforms still publish perfectly good feeds that nobody uses because everyone assumes RSS 301 died. These are keyless, stable, cheap to poll and attach to no account — the best monitoring 302 surface available, where it exists. 303 304 ```bash 305 # YouTube channel, by channel ID (the UC... form; @handles do not work here) 306 curl -s 'https://www.youtube.com/feeds/videos.xml?channel_id=UCXuqSBlHAE6Xw-yeJA0Tunw' 307 308 # a subreddit's new posts 309 curl -s -A 'research-monitor/1.0 (contact: you@example.org)' \ 310 'https://www.reddit.com/r/osint/new/.rss' 311 312 # one Reddit user's activity 313 curl -s -A 'research-monitor/1.0 (contact: you@example.org)' \ 314 'https://www.reddit.com/user/someuser/.rss' 315 316 # any Mastodon account, on any instance 317 curl -s 'https://mastodon.social/@Gargron.rss' 318 319 # a Bluesky profile 320 curl -s 'https://bsky.app/profile/bsky.app/rss' 321 322 # a GitHub user's public activity, which dates account behaviour precisely 323 curl -s 'https://github.com/bellingcat.atom' 324 325 # pull just the timestamps, to see cadence without reading content 326 curl -s 'https://www.youtube.com/feeds/videos.xml?channel_id=UCXuqSBlHAE6Xw-yeJA0Tunw' \ 327 | grep -oE '<published>[^<]+' | sed 's/<published>//' 328 ``` 329 330 Reddit rate-limits these hard and will return HTTP 429 to an anonymous, default-User-Agent client 331 within a handful of requests; a descriptive User-Agent and a gap of several seconds between calls 332 fixes it, and Reddit's `search.rss` endpoint is throttled more aggressively than the subreddit and 333 user feeds. YouTube's feed carries only the most recent entries, so it monitors but does not 334 backfill. Bluesky's profile RSS covers posts and not replies or likes — the AT Protocol endpoints 335 on [Social Media Platforms](/sheets/osint/social-media-platforms) go deeper. And X/Twitter publishes 336 nothing of this kind: `snscrape` has been non-functional against it since the 2023 access changes 337 and the public Nitter instances are largely dead, so there is no quiet monitoring surface for X at 338 all — only a logged-in session, with everything that implies. 339 340 ### RSSHub and RSS-Bridge 341 342 For the platforms that killed their feeds. Both are self-hosted services that scrape a site and 343 re-emit it as RSS or Atom, so your monitor still speaks one protocol no matter how many platforms 344 are behind it. 345 346 ```bash 347 # RSSHub -- the larger of the two, ~1000 routes 348 docker run -d --name rsshub -p 1200:1200 diygod/rsshub 349 # or with the compose file, which adds redis caching and a browser for JS sites 350 wget https://raw.githubusercontent.com/DIYgod/RSSHub/master/docker-compose.yml 351 docker compose up -d 352 353 # routes are paths. a Telegram public channel: 354 curl -s 'http://localhost:1200/telegram/channel/awesomeRSSHub' 355 # the route catalogue lives at docs.rsshub.app/routes/ 356 357 # RSS-Bridge -- fewer routes (~450) but simpler, and a web UI that builds the URL 358 docker create --name=rss-bridge --publish 3000:80 \ 359 --volume $(pwd)/config:/config rssbridge/rss-bridge 360 docker start rss-bridge 361 362 # bridges are query parameters rather than paths 363 curl -s 'http://localhost:3000/?action=display&bridge=<BridgeName>&format=Atom' 364 ``` 365 366 Self-host both. The public `rsshub.app` demo instance returns HTTP 403 to ordinary requests and is 367 not a reliable backend for anything you depend on, and a shared public instance makes your watch 368 list someone else's log file either way. Expect individual routes to break: they are scrapers 369 wearing an RSS hat, and when a platform changes its markup the route returns an empty feed rather 370 than an error — which reads exactly like "the target stopped posting". Check periodically that a 371 route still returns items, and never conclude silence from an empty bridge feed without confirming 372 against the site. 373 374 ### curl and jq on a schedule 375 376 For a public JSON endpoint, a few lines beat any framework. The pattern that matters is the 377 watermark: store the timestamp of the newest item you have seen, and ask only for things after it. 378 379 ```bash 380 # poll Arctic Shift for new Reddit posts in a subreddit since the last run. 381 # full Arctic Shift usage is on the social media platforms sheet; this is the 382 # monitoring shape of it 383 AS='https://arctic-shift.photon-reddit.com/api' 384 STATE=~/monitor/last_seen_osint 385 SINCE=$(cat "$STATE" 2>/dev/null || echo 0) 386 387 curl -sS --fail -G "$AS/posts/search" \ 388 --data-urlencode 'subreddit=osint' \ 389 --data-urlencode "after=$SINCE" \ 390 --data-urlencode 'sort=asc' \ 391 --data-urlencode 'limit=100' \ 392 -o /tmp/new.json || { echo "poll failed $(date -u +%FT%TZ)" >> ~/monitor/errors.log; exit 1; } 393 394 # advance the watermark only on a successful fetch, or a failure silently 395 # becomes a gap you never notice 396 jq -r '.data[-1].created_utc // empty' /tmp/new.json | grep . && \ 397 jq -r '.data[-1].created_utc' /tmp/new.json > "$STATE" 398 399 # append raw, then derive. the raw file is the thing you cannot regenerate 400 cat /tmp/new.json >> ~/monitor/osint-raw.ndjson 401 jq -r '.data[] | [.created_utc, .author, .title] | @tsv' /tmp/new.json 402 ``` 403 404 ```bash 405 # a Bluesky author feed, keyless, as a second example of the same shape 406 curl -sS --fail 'https://public.api.bsky.app/xrpc/app.bsky.feed.getAuthorFeed?actor=bsky.app&limit=50' \ 407 | jq -r '.feed[] | [.post.indexedAt, .post.record.text] | @tsv' 408 409 # crontab: daily, at a minute that is not :00, with failures mailed to you 410 # 37 6 * * * /home/you/monitor/poll-osint.sh >> /home/you/monitor/run.log 2>&1 411 ``` 412 413 `--fail` is the flag that turns a silent HTML error page into a non-zero exit, and without it your 414 monitor will happily append a Cloudflare block page to its dataset for a month. Advance the 415 watermark only after a confirmed good response, log every failure with a timestamp, and treat the 416 log as part of the dataset — the difference between "they went quiet" and "we were blocked" lives 417 there and nowhere else. Note that Bluesky's `getAuthorFeed` is open without a token but 418 `searchPosts` on the same public host returns 403 unauthenticated, which is the sort of asymmetry 419 worth checking per endpoint before you build a schedule around it. 420 421 ### Arctic Shift for Reddit history 422 423 The live successor to Pushshift, and the only practical way to get Reddit content that has since 424 been deleted or edited. For monitoring it plays a specific role: it is the backfill that gives your 425 forward-looking poll a baseline, and the thing you check when a post you captured disappears. 426 427 ```bash 428 AS='https://arctic-shift.photon-reddit.com/api' 429 430 # the baseline: what this account did before you started watching it 431 curl -s -G "$AS/comments/search" --data-urlencode 'author=some_user' \ 432 --data-urlencode 'sort=asc' --data-urlencode 'limit=100' \ 433 | jq -r '.data[] | [.created_utc, .subreddit] | @tsv' 434 435 # posting cadence by hour of day, which is what exposes a coordinated account 436 curl -s -G "$AS/posts/search" --data-urlencode 'author=some_user' \ 437 --data-urlencode 'limit=500' --data-urlencode 'fields=created_utc' \ 438 | jq -r '.data[].created_utc' \ 439 | python3 -c 'import sys,datetime,collections; c=collections.Counter(datetime.datetime.fromtimestamp(int(l), datetime.UTC).hour for l in sys.stdin); [print(f"{h:02d} {c[h]}") for h in range(24)]' 440 441 # did a post you captured actually get removed, or did you lose it 442 curl -s -G "$AS/posts/ids" --data-urlencode 'ids=t3_abc123' | jq '.data[0] | {title, selftext, removed_by_category}' 443 ``` 444 445 Keyword search needs an accompanying `author`, `subreddit`, `link_id` or `parent_id` — bare keyword 446 sweeps are refused, and so are very broad author or subreddit queries. The archive holds what it 447 ingested at the time, so an edit made after ingestion is invisible and a post deleted within 448 seconds may never have been captured at all; a record here proves the text was published, not that 449 it is live. Bulk dumps for offline work are published separately from the API. Query syntax in full 450 is on [Social Media Platforms](/sheets/osint/social-media-platforms). 451 452 ## Narrative and coordination analysis 453 454 - **Posting-time clustering** exposes networks. Accounts that post within seconds of each other, 455 repeatedly, are coordinated. The `jq` plus histogram pattern above is enough to see it. 456 - **Identical phrasing across accounts** is the strongest single indicator of copy-paste campaigns. 457 - **Follower overlap** between accounts is more telling than follower count. 458 - [The Information Laundromat](https://informationlaundromat.com/) compares content and technical 459 metadata — ad IDs, analytics tags, registration details, code structure — across sites, to find 460 where the same text is syndicated and which sites share infrastructure. Built by the Alliance for 461 Securing Democracy, which merged into the Institute for Strategic Dialogue on 1 January 2026; the 462 tool remains live at the address above. The older hyphenated domain no longer resolves. 463 - Graphing and clustering what you have collected is on 464 [Data Analysis & Visualisation](/sheets/osint/data-analysis-and-visualisation). 465 466 ## Tool reference 467 468 | Tool | What it does | Cost | 469 | --- | --- | --- | 470 | [4CAT](https://4cat.nl/) | 4CAT is a tool designed for the easy collection and analysis of online datasets. It allows researchers to uncover patterns and trends in data from social… | free | 471 | [Atlos](https://www.atlos.org/) | ATLOS is a platform for collaborative and large-scale open source investigations. | partly free | 472 | [Datasette](https://datasette.io) | Open-source “WordPress-for-data” that turns any SQLite database into an interactive website and JSON API in seconds; ideal for publishing, exploring and… | free | 473 | [Gephi](https://gephi.org) | Open-source network analysis and visualization software | free | 474 | [Maltego Graph](https://www.maltego.com/downloads/) | Maltego Graph is an investigation platform that combines two things at once: (1) It acts as a search tool, and (2) It creates a graph establishing links… | partly free | 475 | [Pinpoint](https://journaliststudio.google.com/pinpoint/about) | A tool by Google to catalogue uploaded documents and files, providing automated text recogntion, indexing, audiotranscriptions and other (AI-powered)… | free | 476 | [Time.Graphics](https://time.graphics) | A tool for creating, visualizing, and managing timelines online. | partly free | 477 | [Zeeschuimer](https://github.com/digitalmethodsinitiative/zeeschuimer) | Firefox extension that records social media data as you browse it, for import into 4CAT. | free | 478 | [changedetection.io](https://changedetection.io/) | Self-hosted page-change monitoring with CSS/XPath/jq filters and diff history. | free | 479 | [Distill](https://distill.io/) | Hosted page-change monitoring, browser extension plus cloud checks. | partly free | 480 481 ## Pitfalls 482 483 - **Dead tools still appear in tutorials.** `snscrape` and the public Nitter instances do not work 484 against X. Bing's Visual Search API and the rest of the Bing Search family were decommissioned in 485 August 2025. Check a tool is alive before you design a collection around it. 486 - **Rate limits end collections mid-run.** Check for gaps before you analyse; a missing day looks 487 like silence rather than a failure. 488 - **An empty feed is ambiguous.** A broken RSSHub route, a changed selector and a target who 489 stopped posting all look identical downstream. Monitor your monitors. 490 - **Sampling bias reads as a finding.** If a tool only reaches accounts above some follower count, 491 or only what you happened to scroll past, its "network" is an artefact of that cutoff. 492 - **Coordination needs a baseline.** Fans of the same thing post about it at the same time. Compare 493 against normal behaviour for the topic before calling it inauthentic. 494 - **Monitoring is contact.** Scheduled requests are a pattern; a logged-in monitor is an 495 attributable pattern. Decide what that costs before you start, not after. 496 - **Storage and legality.** Bulk personal data attracts data-protection obligations even when every 497 individual item was public, and a dataset is harder to delete than to collect. 498 499 ## Worked example 500 501 One datum: a single Telegram channel name, from a screenshot, alleged to be seeding a story that 502 later appeared on a cluster of news-like websites. 503 504 1. **Baseline before watching.** Pull the channel's existing history into 4CAT as a Telegram 505 dataset with a bounded date range. You now know its normal posting cadence and vocabulary, which 506 is what any later claim of a surge has to be measured against. 507 2. **Find the forward-looking surface.** Telegram has no public feed, so an RSSHub route 508 (`/telegram/channel/<name>` on your own instance) gives you a pollable endpoint. Verify it 509 returns items today, so that an empty feed next month means something. 510 3. **Schedule it honestly.** A daily `curl` with a stored watermark, a descriptive User-Agent, and 511 a failure log. Not hourly — a story that takes days to syndicate does not need fifteen-minute 512 resolution, and the slower poll survives longer. 513 4. **Watch the downstream sites for edits, not just posts.** Each suspected site gets a 514 `changedetection.io` watch filtered to the article body, so a quietly amended paragraph or a 515 removed byline raises an alert with a stored diff. This is the part that a collection-only 516 approach misses entirely. 517 5. **Capture the video claims properly.** The channel posts clips; `yt-dlp --download-archive` 518 with `--write-info-json` on its linked channels picks up only what is new each day and records 519 `upload_date` for each, giving every clip a latest-possible date. 520 6. **Test the syndication claim.** Feed one article URL to the Information Laundromat. Content 521 similarity tells you the text is shared; the technical indicators — a common analytics ID across 522 four of the sites — tell you something stronger, because wording can be copied by anyone and a 523 shared tracking ID usually cannot. 524 7. **Cross-check the Reddit leg.** Arctic Shift for posts linking those domains, grouped by 525 subreddit and by hour, shows whether the amplification is a handful of accounts on a schedule or 526 genuine spread. 527 8. **Archive as you go.** Every page you will cite goes to a snapshot at the time you saw it, per 528 [Archiving & Evidence](/sheets/osint/archiving-and-evidence) — these sites edit and disappear, 529 which is the behaviour you are documenting. 530 531 What you can assert: a dated posting history for the channel, dated first appearances on each 532 downstream site, stored diffs of any subsequent edits, and a shared technical indicator linking 533 some of those sites. 534 535 What would falsify it: an RSSHub route that broke silently mid-period, which would turn a real gap 536 into an apparent one — so the run log, showing a successful fetch every day, is doing as much 537 evidential work as the data. A shared analytics ID that turns out to belong to a common CMS 538 template, or to an agency that serves unrelated clients, would break the infrastructure link; check 539 what else carries that ID before leaning on it. 540 541 ## Broader catalogues 542 543 - [Social Media OSINT](https://tools.osintnewsletter.com/tool-categories/social-media-osint) 544 545 546 ## Sources 547 548 Both catalogues below are maintained by other people and are considerably larger than 549 this page. Use them as the canonical index; this sheet is a working route through them. 550 551 - [Bellingcat's Online Investigation Toolkit](https://bellingcat.gitbook.io/toolkit) — ~340 tools, each with its own 552 review page covering cost, difficulty, requirements and limitations. 553 - [OSINT Newsletter Tools Library](https://tools.osintnewsletter.com) — ~280 tools, organised by investigative goal. 554 555 Neither publishes a licence, so nothing here is copied from them: tool names, one-line 556 descriptions, cost flags and links are catalogue facts, and the method and commentary are 557 this site's own. See [credits](/credits).