daemon-sec-cheatsheet

The cheatsheet vault for operators: AD, enumeration, exploitation, priv-esc, web, DFIR
git clone https://git.daemon-sec.xyz/daemon-sec-cheatsheet.git
Log | Files | Refs | README | LICENSE

social-media-monitoring.md (31354B)


      1 ---
      2 title: "Cross-Platform Monitoring & Collection"
      3 description: "Track accounts, hashtags and narratives across several platforms at once, and watch pages for changes."
      4 category: osint
      5 subcategory: "Social Media"
      6 tags: [osint, monitoring, collection, social-media]
      7 tools: [4cat, zeeschuimer, changedetection.io, distill, yt-dlp, rsshub, rss-bridge, arctic-shift]
      8 difficulty: intermediate
      9 updated: 2026-10-04
     10 references:
     11   - name: "Bellingcat's Online Investigation Toolkit"
     12     url: "https://bellingcat.gitbook.io/toolkit"
     13     author: "Bellingcat"
     14     license: none
     15     relation: derived
     16     note: "Tool catalogue: names, descriptions, cost flags and links for this area."
     17   - name: "OSINT Newsletter Tools Library"
     18     url: "https://tools.osintnewsletter.com"
     19     author: "The OSINT Newsletter"
     20     license: none
     21     relation: derived
     22     note: "Second tool catalogue, cross-checked against the above."
     23 ---
     24 
     25 ## What this covers
     26 
     27 Watching things over time instead of looking once, and holding what you collect so you can analyse
     28 it later without re-collecting. Two distinct jobs with different tooling: **bulk collection** of a
     29 corpus for analysis, and **change detection** on a small number of pages you care about. The third
     30 thing, which is not optional, is doing either without the pattern of your collection becoming
     31 visible to the people you are collecting on.
     32 
     33 ## Method
     34 
     35 1. **Decide what you are watching before you build anything.** A corpus for network analysis and an
     36    alert on one bio edit need completely different infrastructure, and building the first when you
     37    needed the second is the usual waste.
     38 2. **Prefer the platform's own feed.** A native RSS or JSON endpoint is stable, keyless, polite and
     39    does not attribute activity to an account. Reach for a scraper only when no feed exists.
     40 3. **Store raw, analyse later.** Keep the untouched response alongside anything you derive from it.
     41    You will want to re-run the analysis with a different question, and you will not want to
     42    re-collect under a worse rate limit.
     43 4. **Record gaps as gaps.** A failed poll is a hole in your data, not a quiet day. Log the failures
     44    next to the results or your time series will lie to you.
     45 5. **Set the cadence to the question.** Hourly is almost never justified. A daily poll that runs
     46    for a year is more valuable, and far less conspicuous, than a five-minute poll that gets you
     47    blocked in a week.
     48 6. **Baseline before you conclude.** Coordination, surges and silences only mean something against
     49    what normal looks like for that topic.
     50 
     51 The judgement calls: whether you need the content or only the fact that it changed — the second is
     52 far cheaper and far quieter. Whether the collection needs an account at all, because the moment it
     53 does, the collection has an identity and a history. And whether you are allowed to keep what you
     54 are about to collect, which is a question to settle at the start rather than after you have a
     55 database of it.
     56 
     57 ## Collection hygiene
     58 
     59 Monitoring is repeated, scheduled, patterned contact with a target's content. One look is
     60 invisible. The same request every fifteen minutes from one address for six months is a signature,
     61 and on some platforms it is a signature attached to a logged-in account.
     62 
     63 ```text
     64 Account
     65   - never your own, and never one that shares a recovery phone or email with your own
     66   - a research account per investigation where the platform allows it; one burned
     67     account should not cost you the others
     68   - assume every logged-in query is retained and attributable, including search terms
     69   - some platforms notify a user when a profile is viewed; know which before you look
     70 
     71 Rate
     72   - daily is the default. Justify anything faster to yourself in writing
     73   - randomise the interval. A poll at exactly :00 every hour is machine-obvious
     74   - respect the stated limit, and treat a 429 as a signal to back off for hours,
     75     not to retry in a loop
     76   - set a real, honest User-Agent that identifies the project. It gets you unblocked
     77     more often than a spoofed browser string does
     78 
     79 Network
     80   - one address for all of it correlates every collection you run
     81   - a residential or VPN exit is a trade: less rate limiting, more attribution risk
     82     to whoever pays for it
     83   - Tor is blocked by most of these platforms, so it is rarely the answer here
     84 
     85 Footprint
     86   - log every request you make, with timestamp and response code. You need it to
     87     distinguish "they stopped posting" from "we got blocked"
     88   - decide retention up front. Bulk personal data attracts obligations even when
     89     every item was public
     90 ```
     91 
     92 The archiving and provenance side of this — hashes, snapshots, what makes a collected item citable
     93 later — is on [Archiving & Evidence](/sheets/osint/archiving-and-evidence). Per-platform surfaces,
     94 query syntax and the specific scrapers are on
     95 [Social Media Platforms](/sheets/osint/social-media-platforms); this sheet is about running those
     96 things repeatedly rather than once.
     97 
     98 ## Key tools
     99 
    100 ### 4CAT
    101 
    102 A self-hosted capture-and-analysis platform. It is the serious option because collection and
    103 analysis live in the same place: the dataset you capture stays on your machine, and you can re-run
    104 a different analysis over it months later without touching the platform again. That is the thing
    105 no hosted service gives you.
    106 
    107 ```bash
    108 # the documented install is the compose file plus its .env, not a clone
    109 mkdir 4cat && cd 4cat
    110 curl -O https://raw.githubusercontent.com/digitalmethodsinitiative/4cat/master/docker-compose.yml
    111 curl -O https://raw.githubusercontent.com/digitalmethodsinitiative/4cat/master/.env
    112 
    113 # the two settings worth reading before you start it:
    114 #   SERVER_BIND_ADDRESS=127.0.0.1   localhost only (the shipped default)
    115 #   PUBLIC_PORT=80                  the port the web UI lands on
    116 grep -E 'SERVER_BIND_ADDRESS|PUBLIC_PORT|DOCKER_TAG' .env
    117 
    118 docker compose up -d
    119 docker compose logs -f backend      # first start builds indexes; wait for it
    120 
    121 # the one-time admin link is printed in the logs, not emailed
    122 docker compose logs frontend | grep -i 'create a new user\|token'
    123 ```
    124 
    125 Creating a dataset, once it is up:
    126 
    127 ```text
    128 1. http://localhost:80 -> "Create dataset".
    129 2. Pick the data source. Natively it collects 4chan, 8kun, Bluesky, Telegram, Tumblr
    130    and TikTok (from a list of URLs). Everything else arrives as an upload.
    131 3. Set the query and the date range. Date range is the field people leave open and
    132    then wonder why the job runs for a day -- bound it.
    133 4. Submit, then leave it. The backend queues and processes; the dataset appears in
    134    your list with a status. Big Telegram or 4chan pulls take hours.
    135 5. On the finished dataset, run processors rather than exporting immediately:
    136    word frequencies, co-word networks, time series, top posters, image downloads
    137    and clustering. Each produces a child dataset you can chain further processors on.
    138 6. Export the one you want as CSV or NDJSON, and keep the parent dataset. The parent
    139    is the thing you cannot re-create.
    140 ```
    141 
    142 The platform coverage is the live constraint and it changes: X/Twitter, Instagram, LinkedIn,
    143 Threads and Pinterest are no longer collected by 4CAT directly — they come in through Zeeschuimer
    144 below, or as an upload from elsewhere. Datasets are big, and the Docker volumes will fill a disk
    145 without warning, so check free space before a large pull. And a 4CAT capture is a snapshot of what
    146 the platform served at that moment: edits and deletions after the capture are invisible to it,
    147 which is a feature when you want the original and a trap when you assume it reflects the live site.
    148 
    149 ### Zeeschuimer
    150 
    151 The companion extension, and the current answer to "how do I get X, Instagram or TikTok data into
    152 4CAT". It watches the data the platform sends to your browser as you scroll, and keeps it. No API,
    153 no scraping requests of its own — it records what your normal browsing already fetched.
    154 
    155 ```text
    156 1. Firefox only. Signed .xpi builds are on the project's releases page; Chrome is
    157    not supported.
    158 2. Install, then open the extension and enable the platform you are about to browse.
    159    Supported: TikTok, Instagram, Threads, X/Twitter, Pinterest, Gab, Truth Social,
    160    9gag, Imgur, Douyin and RedNote/Xiaohongshu.
    161 3. Browse normally -- a profile, a hashtag, a search. The item counter in the
    162    extension ticks up as content loads. Content you never scrolled past is not
    163    captured, because it was never sent to your browser.
    164 4. Export as NDJSON or CSV, or put your 4CAT URL in the extension and upload straight
    165    into it as a dataset.
    166 5. Record how you browsed: which profile, which tab, how far you scrolled. That is
    167    the sampling frame, and without it the dataset's coverage is undocumented.
    168 ```
    169 
    170 This is manual collection with automatic recording, so it does not scale and it does not run
    171 unattended — which is also why it survives platform changes better than scrapers do. It captures
    172 through your logged-in session, so everything you collect is attributable to that account: use a
    173 research profile in a separate Firefox container or profile, never your own. Platform support needs
    174 constant maintenance and individual platforms do break; check the releases page before assuming the
    175 extension is at fault rather than the site.
    176 
    177 ### changedetection.io
    178 
    179 Self-hosted page watching. This is the one to run when the thing you need to notice is an edit — a
    180 bio rewritten, a post deleted, an officer removed from a company page, a policy document quietly
    181 amended, a price or a staff list changing.
    182 
    183 ```bash
    184 # docker, bound to localhost
    185 docker run -d --restart always -p "127.0.0.1:5000:5000" \
    186   -v datastore-volume:/datastore \
    187   --name changedetection.io dgtlmoon/changedetection.io
    188 
    189 # or as a package, if you would rather not run a container
    190 pip3 install changedetection.io
    191 changedetection.io -d /path/to/empty/data/dir -p 5000
    192 ```
    193 
    194 ```text
    195 1. http://localhost:5000 -> "Add new watch", paste the URL.
    196 2. Open "Edit" and set the filter before you set anything else. CSS selector, XPath,
    197    JSONPath or jq -- point it at the narrowest element carrying the signal. Watching
    198    a whole page means alerting on rotating ads, view counters and timestamps.
    199    The "Visual Selector" tab picks the element by clicking it.
    200 3. Set "Ignore text" for the lines you know churn: relative dates, "N views",
    201    cookie-banner text.
    202 4. Recheck time: per-watch. Daily for most things. The default applies to every watch
    203    you add, so set it low once rather than per watch.
    204 5. Notifications: Discord, email, Slack, Telegram or a webhook via Apprise. A webhook
    205    into your own notes or ticketing is the one that leaves a record.
    206 6. For a page that renders client-side, enable the Playwright/Sockpuppetbrowser
    207    fetcher for that watch -- the default plain fetch sees an empty shell and will
    208    report "no change" forever.
    209 7. "Browser Steps" handles a page behind a form: click, fill, submit, then diff what
    210    comes back.
    211 8. Every change is stored as a snapshot with a diff view, which is the part that makes
    212    it evidence rather than an alert.
    213 ```
    214 
    215 Running it yourself means your watch list is not a third party's business record, which matters
    216 when the watch list itself is sensitive. The cost is that a JS-heavy watch runs a real browser and
    217 is far heavier than a text diff — a dozen of those on a small VPS will struggle. And a watch only
    218 sees what an unauthenticated fetch from your server sees: a page that needs a login, or that
    219 geo-varies, needs Browser Steps or will silently watch the wrong thing.
    220 
    221 ### Distill
    222 
    223 The hosted equivalent, for when you will not run infrastructure. Same idea, less setup, and a free
    224 tier that is genuinely usable for a handful of watches: **25 monitors total but only 5 in the
    225 cloud, a 6-hour minimum cloud interval, 1,000 cloud checks a month, 2 devices, and 30 email alerts
    226 a month**. Local monitors in the browser extension are unlimited but only run while the browser is
    227 open.
    228 
    229 ```text
    230 1. Browser extension or distill.io. On the page you want, click the extension and
    231    select the region -- it generates the selector for you.
    232 2. Choose local (runs in your browser, unlimited, only while open) or cloud (runs
    233    without you, capped as above). For anything that matters, cloud.
    234 3. Set the check interval. Anything under 6 hours is a paid feature on the free tier,
    235    and six-hourly is adequate for almost all of this work anyway.
    236 4. Set the condition, not just "any change" -- Distill supports text conditions, so
    237    "alert when the number changes" beats "alert when the page differs".
    238 5. Alerts to email or webhook; the email allowance is the binding constraint on free.
    239 6. Export the watch list as JSON periodically. It is the only part that is painful to
    240    rebuild.
    241 ```
    242 
    243 Your watch list and every snapshot live on their servers, which is the trade for not running
    244 anything. For a target that could plausibly subpoena or compromise a third party, that is the wrong
    245 trade and `changedetection.io` is the answer instead. The free tier's 6-hour floor also means you
    246 will miss a post that goes up and comes down inside a window, which is precisely the kind of
    247 deletion worth catching — if that is the scenario, self-host and poll faster.
    248 
    249 ### yt-dlp with a download archive
    250 
    251 For recurring capture of a channel's output, the archive file is the whole trick: it records the ID
    252 of everything already fetched, so the next run picks up only what is new. That turns a one-off
    253 download into a monitor you can cron.
    254 
    255 ```bash
    256 pipx install yt-dlp
    257 
    258 # first run: establish the archive. --break-on-existing stops as soon as it meets
    259 # something already recorded, so later runs walk only the new items
    260 yt-dlp --download-archive archive.txt --break-on-existing --lazy-playlist \
    261   -o '%(upload_date)s-%(id)s.%(ext)s' \
    262   'https://youtube.com/@channel/videos'
    263 
    264 # metadata-only monitoring: no video files, just the record that it existed
    265 yt-dlp --download-archive seen.txt --break-on-existing \
    266   --skip-download --write-info-json --write-thumbnail \
    267   'https://youtube.com/@channel/videos'
    268 
    269 # bound it by date instead, for a backfill of a known window
    270 yt-dlp --dateafter 20260101 --datebefore 20260401 --download-archive archive.txt URL
    271 
    272 # subtitles, which turn a channel's output into greppable text as it arrives
    273 yt-dlp --download-archive seen.txt --break-on-existing --skip-download \
    274   --write-auto-subs --sub-langs en 'https://youtube.com/@channel/videos'
    275 
    276 # throttle it. these two flags are the difference between a monitor and a nuisance
    277 yt-dlp --sleep-requests 2 --sleep-interval 10 --max-sleep-interval 30 \
    278   --download-archive archive.txt URL
    279 
    280 # several channels in one run, each stopping at its own first-seen item
    281 yt-dlp --break-per-input --break-on-existing --download-archive archive.txt \
    282   -a channels.txt
    283 
    284 # a cap, so a misconfigured run cannot pull a thousand files overnight
    285 yt-dlp --max-downloads 50 --download-archive archive.txt URL
    286 ```
    287 
    288 The archive file is state: back it up, and never delete it to "start fresh" unless you mean to
    289 re-download everything. `--break-on-existing` assumes the listing is newest-first, which it is for
    290 channels and playlists and is not for some search result pages — on those, drop it and let the
    291 archive do the skipping. Extractors break when platforms change their pages, so `pipx upgrade
    292 yt-dlp` before blaming a URL, and a cron job that has silently failed for three weeks is worse than
    293 no monitoring at all, so alert on non-zero exits. Adding `--cookies-from-browser` attaches a real
    294 session to every request and makes the whole monitor attributable; avoid it unless the content
    295 genuinely requires a login. Full metadata usage is on
    296 [Social Media Platforms](/sheets/osint/social-media-platforms).
    297 
    298 ### Native feeds, before you reach for a bridge
    299 
    300 Several platforms still publish perfectly good feeds that nobody uses because everyone assumes RSS
    301 died. These are keyless, stable, cheap to poll and attach to no account — the best monitoring
    302 surface available, where it exists.
    303 
    304 ```bash
    305 # YouTube channel, by channel ID (the UC... form; @handles do not work here)
    306 curl -s 'https://www.youtube.com/feeds/videos.xml?channel_id=UCXuqSBlHAE6Xw-yeJA0Tunw'
    307 
    308 # a subreddit's new posts
    309 curl -s -A 'research-monitor/1.0 (contact: you@example.org)' \
    310   'https://www.reddit.com/r/osint/new/.rss'
    311 
    312 # one Reddit user's activity
    313 curl -s -A 'research-monitor/1.0 (contact: you@example.org)' \
    314   'https://www.reddit.com/user/someuser/.rss'
    315 
    316 # any Mastodon account, on any instance
    317 curl -s 'https://mastodon.social/@Gargron.rss'
    318 
    319 # a Bluesky profile
    320 curl -s 'https://bsky.app/profile/bsky.app/rss'
    321 
    322 # a GitHub user's public activity, which dates account behaviour precisely
    323 curl -s 'https://github.com/bellingcat.atom'
    324 
    325 # pull just the timestamps, to see cadence without reading content
    326 curl -s 'https://www.youtube.com/feeds/videos.xml?channel_id=UCXuqSBlHAE6Xw-yeJA0Tunw' \
    327   | grep -oE '<published>[^<]+' | sed 's/<published>//'
    328 ```
    329 
    330 Reddit rate-limits these hard and will return HTTP 429 to an anonymous, default-User-Agent client
    331 within a handful of requests; a descriptive User-Agent and a gap of several seconds between calls
    332 fixes it, and Reddit's `search.rss` endpoint is throttled more aggressively than the subreddit and
    333 user feeds. YouTube's feed carries only the most recent entries, so it monitors but does not
    334 backfill. Bluesky's profile RSS covers posts and not replies or likes — the AT Protocol endpoints
    335 on [Social Media Platforms](/sheets/osint/social-media-platforms) go deeper. And X/Twitter publishes
    336 nothing of this kind: `snscrape` has been non-functional against it since the 2023 access changes
    337 and the public Nitter instances are largely dead, so there is no quiet monitoring surface for X at
    338 all — only a logged-in session, with everything that implies.
    339 
    340 ### RSSHub and RSS-Bridge
    341 
    342 For the platforms that killed their feeds. Both are self-hosted services that scrape a site and
    343 re-emit it as RSS or Atom, so your monitor still speaks one protocol no matter how many platforms
    344 are behind it.
    345 
    346 ```bash
    347 # RSSHub -- the larger of the two, ~1000 routes
    348 docker run -d --name rsshub -p 1200:1200 diygod/rsshub
    349 # or with the compose file, which adds redis caching and a browser for JS sites
    350 wget https://raw.githubusercontent.com/DIYgod/RSSHub/master/docker-compose.yml
    351 docker compose up -d
    352 
    353 # routes are paths. a Telegram public channel:
    354 curl -s 'http://localhost:1200/telegram/channel/awesomeRSSHub'
    355 # the route catalogue lives at docs.rsshub.app/routes/
    356 
    357 # RSS-Bridge -- fewer routes (~450) but simpler, and a web UI that builds the URL
    358 docker create --name=rss-bridge --publish 3000:80 \
    359   --volume $(pwd)/config:/config rssbridge/rss-bridge
    360 docker start rss-bridge
    361 
    362 # bridges are query parameters rather than paths
    363 curl -s 'http://localhost:3000/?action=display&bridge=<BridgeName>&format=Atom'
    364 ```
    365 
    366 Self-host both. The public `rsshub.app` demo instance returns HTTP 403 to ordinary requests and is
    367 not a reliable backend for anything you depend on, and a shared public instance makes your watch
    368 list someone else's log file either way. Expect individual routes to break: they are scrapers
    369 wearing an RSS hat, and when a platform changes its markup the route returns an empty feed rather
    370 than an error — which reads exactly like "the target stopped posting". Check periodically that a
    371 route still returns items, and never conclude silence from an empty bridge feed without confirming
    372 against the site.
    373 
    374 ### curl and jq on a schedule
    375 
    376 For a public JSON endpoint, a few lines beat any framework. The pattern that matters is the
    377 watermark: store the timestamp of the newest item you have seen, and ask only for things after it.
    378 
    379 ```bash
    380 # poll Arctic Shift for new Reddit posts in a subreddit since the last run.
    381 # full Arctic Shift usage is on the social media platforms sheet; this is the
    382 # monitoring shape of it
    383 AS='https://arctic-shift.photon-reddit.com/api'
    384 STATE=~/monitor/last_seen_osint
    385 SINCE=$(cat "$STATE" 2>/dev/null || echo 0)
    386 
    387 curl -sS --fail -G "$AS/posts/search" \
    388   --data-urlencode 'subreddit=osint' \
    389   --data-urlencode "after=$SINCE" \
    390   --data-urlencode 'sort=asc' \
    391   --data-urlencode 'limit=100' \
    392   -o /tmp/new.json || { echo "poll failed $(date -u +%FT%TZ)" >> ~/monitor/errors.log; exit 1; }
    393 
    394 # advance the watermark only on a successful fetch, or a failure silently
    395 # becomes a gap you never notice
    396 jq -r '.data[-1].created_utc // empty' /tmp/new.json | grep . && \
    397   jq -r '.data[-1].created_utc' /tmp/new.json > "$STATE"
    398 
    399 # append raw, then derive. the raw file is the thing you cannot regenerate
    400 cat /tmp/new.json >> ~/monitor/osint-raw.ndjson
    401 jq -r '.data[] | [.created_utc, .author, .title] | @tsv' /tmp/new.json
    402 ```
    403 
    404 ```bash
    405 # a Bluesky author feed, keyless, as a second example of the same shape
    406 curl -sS --fail 'https://public.api.bsky.app/xrpc/app.bsky.feed.getAuthorFeed?actor=bsky.app&limit=50' \
    407   | jq -r '.feed[] | [.post.indexedAt, .post.record.text] | @tsv'
    408 
    409 # crontab: daily, at a minute that is not :00, with failures mailed to you
    410 # 37 6 * * *  /home/you/monitor/poll-osint.sh >> /home/you/monitor/run.log 2>&1
    411 ```
    412 
    413 `--fail` is the flag that turns a silent HTML error page into a non-zero exit, and without it your
    414 monitor will happily append a Cloudflare block page to its dataset for a month. Advance the
    415 watermark only after a confirmed good response, log every failure with a timestamp, and treat the
    416 log as part of the dataset — the difference between "they went quiet" and "we were blocked" lives
    417 there and nowhere else. Note that Bluesky's `getAuthorFeed` is open without a token but
    418 `searchPosts` on the same public host returns 403 unauthenticated, which is the sort of asymmetry
    419 worth checking per endpoint before you build a schedule around it.
    420 
    421 ### Arctic Shift for Reddit history
    422 
    423 The live successor to Pushshift, and the only practical way to get Reddit content that has since
    424 been deleted or edited. For monitoring it plays a specific role: it is the backfill that gives your
    425 forward-looking poll a baseline, and the thing you check when a post you captured disappears.
    426 
    427 ```bash
    428 AS='https://arctic-shift.photon-reddit.com/api'
    429 
    430 # the baseline: what this account did before you started watching it
    431 curl -s -G "$AS/comments/search" --data-urlencode 'author=some_user' \
    432   --data-urlencode 'sort=asc' --data-urlencode 'limit=100' \
    433   | jq -r '.data[] | [.created_utc, .subreddit] | @tsv'
    434 
    435 # posting cadence by hour of day, which is what exposes a coordinated account
    436 curl -s -G "$AS/posts/search" --data-urlencode 'author=some_user' \
    437   --data-urlencode 'limit=500' --data-urlencode 'fields=created_utc' \
    438   | jq -r '.data[].created_utc' \
    439   | python3 -c 'import sys,datetime,collections; c=collections.Counter(datetime.datetime.fromtimestamp(int(l), datetime.UTC).hour for l in sys.stdin); [print(f"{h:02d} {c[h]}") for h in range(24)]'
    440 
    441 # did a post you captured actually get removed, or did you lose it
    442 curl -s -G "$AS/posts/ids" --data-urlencode 'ids=t3_abc123' | jq '.data[0] | {title, selftext, removed_by_category}'
    443 ```
    444 
    445 Keyword search needs an accompanying `author`, `subreddit`, `link_id` or `parent_id` — bare keyword
    446 sweeps are refused, and so are very broad author or subreddit queries. The archive holds what it
    447 ingested at the time, so an edit made after ingestion is invisible and a post deleted within
    448 seconds may never have been captured at all; a record here proves the text was published, not that
    449 it is live. Bulk dumps for offline work are published separately from the API. Query syntax in full
    450 is on [Social Media Platforms](/sheets/osint/social-media-platforms).
    451 
    452 ## Narrative and coordination analysis
    453 
    454 - **Posting-time clustering** exposes networks. Accounts that post within seconds of each other,
    455   repeatedly, are coordinated. The `jq` plus histogram pattern above is enough to see it.
    456 - **Identical phrasing across accounts** is the strongest single indicator of copy-paste campaigns.
    457 - **Follower overlap** between accounts is more telling than follower count.
    458 - [The Information Laundromat](https://informationlaundromat.com/) compares content and technical
    459   metadata — ad IDs, analytics tags, registration details, code structure — across sites, to find
    460   where the same text is syndicated and which sites share infrastructure. Built by the Alliance for
    461   Securing Democracy, which merged into the Institute for Strategic Dialogue on 1 January 2026; the
    462   tool remains live at the address above. The older hyphenated domain no longer resolves.
    463 - Graphing and clustering what you have collected is on
    464   [Data Analysis & Visualisation](/sheets/osint/data-analysis-and-visualisation).
    465 
    466 ## Tool reference
    467 
    468 | Tool | What it does | Cost |
    469 | --- | --- | --- |
    470 | [4CAT](https://4cat.nl/) | 4CAT is a tool designed for the easy collection and analysis of online datasets. It allows researchers to uncover patterns and trends in data from social… | free |
    471 | [Atlos](https://www.atlos.org/) | ATLOS is a platform for collaborative and large-scale open source investigations. | partly free |
    472 | [Datasette](https://datasette.io) | Open-source “WordPress-for-data” that turns any SQLite database into an interactive website and JSON API in seconds; ideal for publishing, exploring and… | free |
    473 | [Gephi](https://gephi.org) | Open-source network analysis and visualization software | free |
    474 | [Maltego Graph](https://www.maltego.com/downloads/) | Maltego Graph is an investigation platform that combines two things at once: (1) It acts as a search tool, and (2) It creates a graph establishing links… | partly free |
    475 | [Pinpoint](https://journaliststudio.google.com/pinpoint/about) | A tool by Google to catalogue uploaded documents and files, providing automated text recogntion, indexing, audiotranscriptions and other (AI-powered)… | free |
    476 | [Time.Graphics](https://time.graphics) | A tool for creating, visualizing, and managing timelines online. | partly free |
    477 | [Zeeschuimer](https://github.com/digitalmethodsinitiative/zeeschuimer) | Firefox extension that records social media data as you browse it, for import into 4CAT. | free |
    478 | [changedetection.io](https://changedetection.io/) | Self-hosted page-change monitoring with CSS/XPath/jq filters and diff history. | free |
    479 | [Distill](https://distill.io/) | Hosted page-change monitoring, browser extension plus cloud checks. | partly free |
    480 
    481 ## Pitfalls
    482 
    483 - **Dead tools still appear in tutorials.** `snscrape` and the public Nitter instances do not work
    484   against X. Bing's Visual Search API and the rest of the Bing Search family were decommissioned in
    485   August 2025. Check a tool is alive before you design a collection around it.
    486 - **Rate limits end collections mid-run.** Check for gaps before you analyse; a missing day looks
    487   like silence rather than a failure.
    488 - **An empty feed is ambiguous.** A broken RSSHub route, a changed selector and a target who
    489   stopped posting all look identical downstream. Monitor your monitors.
    490 - **Sampling bias reads as a finding.** If a tool only reaches accounts above some follower count,
    491   or only what you happened to scroll past, its "network" is an artefact of that cutoff.
    492 - **Coordination needs a baseline.** Fans of the same thing post about it at the same time. Compare
    493   against normal behaviour for the topic before calling it inauthentic.
    494 - **Monitoring is contact.** Scheduled requests are a pattern; a logged-in monitor is an
    495   attributable pattern. Decide what that costs before you start, not after.
    496 - **Storage and legality.** Bulk personal data attracts data-protection obligations even when every
    497   individual item was public, and a dataset is harder to delete than to collect.
    498 
    499 ## Worked example
    500 
    501 One datum: a single Telegram channel name, from a screenshot, alleged to be seeding a story that
    502 later appeared on a cluster of news-like websites.
    503 
    504 1. **Baseline before watching.** Pull the channel's existing history into 4CAT as a Telegram
    505    dataset with a bounded date range. You now know its normal posting cadence and vocabulary, which
    506    is what any later claim of a surge has to be measured against.
    507 2. **Find the forward-looking surface.** Telegram has no public feed, so an RSSHub route
    508    (`/telegram/channel/<name>` on your own instance) gives you a pollable endpoint. Verify it
    509    returns items today, so that an empty feed next month means something.
    510 3. **Schedule it honestly.** A daily `curl` with a stored watermark, a descriptive User-Agent, and
    511    a failure log. Not hourly — a story that takes days to syndicate does not need fifteen-minute
    512    resolution, and the slower poll survives longer.
    513 4. **Watch the downstream sites for edits, not just posts.** Each suspected site gets a
    514    `changedetection.io` watch filtered to the article body, so a quietly amended paragraph or a
    515    removed byline raises an alert with a stored diff. This is the part that a collection-only
    516    approach misses entirely.
    517 5. **Capture the video claims properly.** The channel posts clips; `yt-dlp --download-archive`
    518    with `--write-info-json` on its linked channels picks up only what is new each day and records
    519    `upload_date` for each, giving every clip a latest-possible date.
    520 6. **Test the syndication claim.** Feed one article URL to the Information Laundromat. Content
    521    similarity tells you the text is shared; the technical indicators — a common analytics ID across
    522    four of the sites — tell you something stronger, because wording can be copied by anyone and a
    523    shared tracking ID usually cannot.
    524 7. **Cross-check the Reddit leg.** Arctic Shift for posts linking those domains, grouped by
    525    subreddit and by hour, shows whether the amplification is a handful of accounts on a schedule or
    526    genuine spread.
    527 8. **Archive as you go.** Every page you will cite goes to a snapshot at the time you saw it, per
    528    [Archiving & Evidence](/sheets/osint/archiving-and-evidence) — these sites edit and disappear,
    529    which is the behaviour you are documenting.
    530 
    531 What you can assert: a dated posting history for the channel, dated first appearances on each
    532 downstream site, stored diffs of any subsequent edits, and a shared technical indicator linking
    533 some of those sites.
    534 
    535 What would falsify it: an RSSHub route that broke silently mid-period, which would turn a real gap
    536 into an apparent one — so the run log, showing a successful fetch every day, is doing as much
    537 evidential work as the data. A shared analytics ID that turns out to belong to a common CMS
    538 template, or to an agency that serves unrelated clients, would break the infrastructure link; check
    539 what else carries that ID before leaning on it.
    540 
    541 ## Broader catalogues
    542 
    543 - [Social Media OSINT](https://tools.osintnewsletter.com/tool-categories/social-media-osint)
    544 
    545 
    546 ## Sources
    547 
    548 Both catalogues below are maintained by other people and are considerably larger than
    549 this page. Use them as the canonical index; this sheet is a working route through them.
    550 
    551 - [Bellingcat's Online Investigation Toolkit](https://bellingcat.gitbook.io/toolkit) — ~340 tools, each with its own
    552   review page covering cost, difficulty, requirements and limitations.
    553 - [OSINT Newsletter Tools Library](https://tools.osintnewsletter.com) — ~280 tools, organised by investigative goal.
    554 
    555 Neither publishes a licence, so nothing here is copied from them: tool names, one-line
    556 descriptions, cost flags and links are catalogue facts, and the method and commentary are
    557 this site's own. See [credits](/credits).