websites-and-infrastructure.md (37658B)
1 --- 2 title: "Websites, Domains & Hosting Infrastructure" 3 description: "Attribute a site: registration history, DNS, certificates, analytics IDs and the hosting it shares with others." 4 category: osint 5 subcategory: "Infrastructure" 6 tags: [osint, domains, dns, certificates, attribution] 7 tools: [urlscan, crtsh, shodan, subfinder, httpx, wayback, rdap] 8 difficulty: intermediate 9 updated: 2026-10-04 10 references: 11 - name: "Bellingcat's Online Investigation Toolkit" 12 url: "https://bellingcat.gitbook.io/toolkit" 13 author: "Bellingcat" 14 license: none 15 relation: derived 16 note: "Tool catalogue: names, descriptions, cost flags and links for this area." 17 - name: "OSINT Newsletter Tools Library" 18 url: "https://tools.osintnewsletter.com" 19 author: "The OSINT Newsletter" 20 license: none 21 relation: derived 22 note: "Second tool catalogue, cross-checked against the above." 23 - name: "Subfinder documentation" 24 url: "https://docs.projectdiscovery.io/opensource/subfinder/usage" 25 author: "ProjectDiscovery" 26 relation: link-only 27 note: "CLI flags and provider-rate-limit syntax were checked against the current usage reference." 28 - name: "httpx documentation" 29 url: "https://docs.projectdiscovery.io/opensource/httpx/usage" 30 author: "ProjectDiscovery" 31 relation: link-only 32 note: "Probe, fingerprint, output and rate-limit flags were checked against the current usage reference." 33 - name: "urlscan API documentation" 34 url: "https://urlscan.io/docs/api/" 35 author: "urlscan GmbH" 36 relation: link-only 37 note: "Search, submission, result and quota endpoint shapes were checked against this reference." 38 --- 39 40 ## What this covers 41 42 Working out who runs a website and what else they run. Passive attribution — historical records, 43 certificate logs, archives — gets you most of the way without ever touching the target. This 44 overlaps with [Enumeration](/enumeration), but the goal here is attribution rather than attack 45 surface. 46 47 ## Method 48 49 1. **Historical WHOIS first.** Current records are almost always privacy-shielded; records from 50 before the shield often are not. 51 2. **Passive DNS** shows which IPs the domain used over time, and which other domains used those 52 IPs. 53 3. **Certificate transparency logs** enumerate subdomains for free, without touching the target, 54 and often expose staging and internal hostnames. 55 4. **Archives** show what the site used to say, including contact details and staff pages since 56 removed. 57 5. **Fingerprint the page.** Analytics IDs, ad IDs, favicon hashes and CMS quirks link sites that 58 share an operator. 59 6. **Check what shares the host.** Shared hosting is meaningless; a dedicated IP hosting five 60 related domains is not. 61 62 ## Page fingerprinting 63 64 Linking two sites to one operator comes down to finding a value that is identical in both and that 65 nobody would share by accident. The signals are not equal: 66 67 | Signal | Why it links sites | Strength | 68 | --- | --- | --- | 69 | Google Analytics / Tag Manager ID | Operators reuse one property across their sites constantly. | Near-conclusive | 70 | AdSense publisher ID | Same, and tied to a payment account. | Near-conclusive | 71 | TLS certificate serial or key reuse | The same certificate on two hosts means one operator. | Strong | 72 | Favicon hash | Shodan indexes it, so one favicon finds every host serving it. | Strong if distinctive | 73 | Distinctive HTML comment or typo | Copy-pasted templates carry unique strings. | Strong | 74 | Response-body hash | Identical pages on different hosts. | Moderate | 75 | Shared IP on dedicated hosting | Meaningful only on a small host, never on a CDN. | Weak to worthless | 76 | Same registrar or nameserver | Millions of unrelated domains share these. | Worthless alone | 77 78 Searching an ID string in [PublicWWW](https://publicwww.com/) or [Grep.app](https://grep.app/) 79 finds the other pages that embed it; `urlscan.io` finds the sites whose scans *requested* it, which 80 catches IDs injected at runtime rather than written into the HTML. 81 82 ## Key tools 83 84 ### RDAP and whois 85 86 RDAP is the structured replacement for `whois`, and the one to reach for first: JSON instead of 87 free text, consistent field names across registries, and no scraping. `whois` is still worth 88 running because some registries put things in the free-text blob that never make it into RDAP. 89 90 ```bash 91 # RDAP, following the bootstrap redirect to whichever registry is authoritative 92 curl -s -L 'https://rdap.org/domain/example.com' | jq '{ldhName,status,events,entities}' 93 94 # the registration and expiry dates on their own 95 curl -s -L 'https://rdap.org/domain/example.com' \ 96 | jq -r '.events[] | "\(.eventAction) \(.eventDate)"' 97 98 # nameserver history is not here, but the current delegation is 99 curl -s -L 'https://rdap.org/domain/example.com' | jq -r '.nameservers[].ldhName' 100 101 # straight to the registry when you know it, which is faster and never rate-limited by a proxy 102 curl -s 'https://rdap.verisign.com/com/v1/domain/example.com' | jq '.events' 103 104 # an IP or a netblock: who holds the allocation, and when it was assigned 105 curl -s -L 'https://rdap.org/ip/93.184.216.34' | jq '{handle,name,country,events}' 106 107 # an ASN, for working out whose network a host actually sits on 108 curl -s -L 'https://rdap.org/autnum/15169' | jq '{handle,name,entities}' 109 110 # the old tool, for the fields RDAP drops 111 whois example.com | grep -iE 'registrar|created|updated|expir|name server|status' 112 ``` 113 114 Current records are privacy-shielded almost everywhere, so RDAP mostly gives you dates, status 115 codes and the registrar — all useful, none of them a name. The creation date is the one field 116 worth trusting: it is set once and does not change, which makes it the backbone of any timeline. 117 Registrar-level and ccTLD RDAP coverage is patchy; a 404 from `rdap.org` means no server is 118 registered for that zone, not that the domain is free. 119 120 ### dig and passive DNS 121 122 Live DNS tells you the present; passive DNS — a third party's historical record of what resolved 123 to what — tells you the past, which is where attribution lives. Run both: the live records name 124 the vendors in use, and the history names the host before the CDN went up. 125 126 ```bash 127 # the basics, worth doing as a set rather than one at a time 128 for r in A AAAA MX NS TXT SOA CNAME; do echo "== $r"; dig +short example.com "$r"; done 129 130 # mail and verification records enumerate the SaaS vendors the operator uses 131 dig +short TXT example.com 132 dig +short TXT _dmarc.example.com 133 dig +short TXT _domainkey.example.com 134 135 # ask the authoritative server directly, bypassing your resolver's cache 136 dig @"$(dig +short NS example.com | head -1)" example.com ANY 137 138 # reverse lookup on the IP, which occasionally names the hosting account 139 dig +short -x 93.184.216.34 140 141 # the full delegation chain, for spotting a subdomain handed to a third party 142 dig +trace staging.example.com | tail -20 143 144 # a zone transfer, which should fail — when it does not, you have the whole zone 145 dig @ns1.example.com example.com AXFR 146 ``` 147 148 `dig +short NS` plus `AXFR` is the one active step here; everything else above only queries 149 resolvers. For history, [DNS History](https://dnshistory.org/) and [Whoxy](https://www.whoxy.com/) 150 are the usual free-tier sources, and [SecurityTrails](https://securitytrails.com/) has the deepest 151 free passive-DNS allowance, metered per month. A shared IP in the history is worthless; a 152 dedicated IP that three domains shared for two years is a finding. 153 154 ### crt.sh 155 156 Certificate transparency logs are public, complete and free, which makes them the best subdomain 157 enumerator in existence — and it never touches the target. crt.sh is the searchable front end, 158 and its JSON output is the part to automate. 159 160 ```bash 161 # every name ever certificated under the domain, deduplicated 162 curl -s 'https://crt.sh/?q=%25.example.com&output=json' \ 163 | jq -r '.[].name_value' | tr ' ' '\n' | sed 's/^\*\.//' | sort -u 164 165 # with the issuance timeline, which dates when each host appeared 166 curl -s 'https://crt.sh/?q=example.com&output=json' \ 167 | jq -r '.[] | "\(.not_before) \(.name_value)"' | sort -u | head -40 168 169 # who issues their certificates — a change of CA often marks a change of operator 170 curl -s 'https://crt.sh/?q=%25.example.com&output=json' \ 171 | jq -r '.[].issuer_name' | sort | uniq -c | sort -rn 172 173 # search by organisation name instead of domain, to find their other estates 174 curl -s 'https://crt.sh/?q=Example+Organisation+Ltd&output=json' | jq -r '.[].common_name' | sort -u 175 176 # one certificate's full detail by crt.sh id 177 curl -s 'https://crt.sh/?id=29510272303&output=json' | jq '{issuer_name,not_before,not_after,serial_number}' 178 179 # only names that still resolve, so you can tell live infrastructure from history 180 curl -s 'https://crt.sh/?q=%25.example.com&output=json' | jq -r '.[].name_value' \ 181 | tr ' ' '\n' | sed 's/^\*\.//' | sort -u \ 182 | while read -r h; do [ -n "$(dig +short "$h")" ] && echo "$h"; done 183 ``` 184 185 crt.sh is run by Sectigo as a public service and is frequently overloaded — a 502 on the web UI is 186 normal and usually clears, and the JSON endpoint stays up more often than the HTML one does. When 187 it is down, `subfinder` pulls the same logs through other indexes. A certificate proves a name was 188 requested, not that a host ever answered on it: staging and internal hostnames show up here 189 precisely because nobody expected them to be published. 190 191 ### subfinder 192 193 Passive subdomain enumeration across several dozen indexes at once — certificate logs, passive 194 DNS providers, search engines, threat-intel feeds. Faster and broader than querying crt.sh alone, 195 and it needs no keys to be useful, though adding free ones roughly doubles what it finds. 196 197 ```bash 198 # install: a single Go binary 199 go install -v github.com/projectdiscovery/subfinder/v2/cmd/subfinder@latest 200 201 # the default run: passive only, every keyless source 202 subfinder -d example.com -silent 203 204 # which sources exist, and which ones your keys unlock 205 subfinder -ls 206 207 # all sources including the slow ones, written to a file 208 subfinder -d example.com -all -o subs.txt 209 210 # JSON lines with the source recorded per host, so a finding is attributable 211 subfinder -d example.com -oJ -cs -o subs.jsonl 212 213 # several root domains in one run 214 subfinder -dL roots.txt -oD ./out -silent 215 216 # only hosts that actually resolve, with their IPs 217 subfinder -d example.com -nW -oI -silent 218 219 # throttle a chatty provider rather than getting your free key banned 220 subfinder -d example.com -rls 'crtsh=1/s' -rl 10 -silent 221 ``` 222 223 Free API keys go in the provider config under your platform's config directory; `subfinder -ls` 224 shows which sources are active. Passive results are historical, so a long list includes hosts 225 retired years ago — pass `-nW` when you need the live set. Sources disagree and some return 226 wildcard noise, so a single-source hit deserves `-cs` output and a second look before you build on 227 it. 228 229 ### urlscan.io 230 231 A hosted browser that visits a page for you, records every request it made, and keeps the result 232 in a searchable archive. That gives you two separate capabilities: look at a hostile page without 233 touching it yourself, and search what *other* people's scans loaded to find sites sharing an 234 unusual third-party request. 235 236 ```bash 237 # search the archive without visiting anything — no key needed for modest volume 238 curl -s 'https://urlscan.io/api/v1/search/?q=domain%3Aexample.com&size=20' \ 239 | jq -r '.results[] | "\(.task.time) \(.page.url)"' 240 241 # every scan whose requested URLs carried a specific analytics ID 242 curl -s -G 'https://urlscan.io/api/v1/search/' \ 243 --data-urlencode 'q=filename:"UA-12345678"' --data-urlencode 'size=100' | jq '.total' 244 245 # scans by the IP a page resolved to, for finding co-hosted estates 246 curl -s 'https://urlscan.io/api/v1/search/?q=page.ip%3A93.184.216.34&size=50' \ 247 | jq -r '.results[].page.domain' | sort -u 248 249 # ASN-wide search, which is how you sweep a bulletproof host 250 curl -s 'https://urlscan.io/api/v1/search/?q=page.asn%3AAS15169&size=100' | jq '.total' 251 252 # submit a scan — the header name is API-Key, and nothing else works 253 curl -s -X POST 'https://urlscan.io/api/v1/scan/' \ 254 -H "API-Key: $URLSCAN_KEY" -H 'Content-Type: application/json' \ 255 -d '{"url":"https://example.com","visibility":"unlisted"}' | jq -r '.uuid' 256 257 # then read the result a minute later: requests, certificates, redirect chain, screenshot 258 curl -s 'https://urlscan.io/api/v1/result/UUID/' \ 259 | jq '{page:.page, redirects:[.data.requests[].request.redirectResponse.url]}' 260 261 # your remaining quota, which is the only reliable statement of your limits 262 curl -s 'https://urlscan.io/user/quotas/' -H "API-Key: $URLSCAN_KEY" 263 ``` 264 265 `visibility` is the decision that matters. A `public` scan is visible to everyone immediately, 266 including your target, who may well be watching the archive for their own domains; `unlisted` 267 keeps it out of search but still reachable by URL. Unauthenticated search quotas are small and 268 reset on the full minute and hour, so batch your queries. A scan is one browser, from a datacentre 269 IP, in one country — a page that cloaks on geography or user agent will show you something the 270 target's visitors never see. 271 272 ### Shodan 273 274 The index of what is listening on the internet, and the only tool here that can answer "what else 275 serves this exact favicon". Its value in attribution work is the pivot: a hash, a certificate 276 serial or an HTTP header becomes a list of hosts. 277 278 ```bash 279 # install and register the key once 280 pipx install shodan 281 shodan init YOUR_API_KEY 282 283 # free-tier reality check: how many results a query has, before spending credits on it 284 shodan count 'http.favicon.hash:-247388890' 285 286 # the pivot that matters — every host serving the same favicon 287 shodan search --fields ip_str,port,hostnames,org 'http.favicon.hash:-247388890' 288 289 # everything known about one address, including historical ports 290 shodan host 93.184.216.34 291 292 # hosts sharing a TLS certificate serial, i.e. the same operator 293 shodan search --fields ip_str,hostnames 'ssl.cert.serial:"73d94ae1858e6f4613fd389a8efe5267"' 294 295 # a distinctive response header or body string, which copy-pasted templates carry 296 shodan search --fields ip_str,http.title 'http.html:"Example Org internal portal"' 297 298 # facet counts, for describing an estate rather than listing it 299 shodan stats --facets country:10,org:10 'ssl.cert.subject.CN:*.example.com' 300 301 # the subdomains and DNS records Shodan has seen for a domain 302 shodan domain example.com 303 304 # bulk download then parse offline, so you pay for the query once 305 shodan download results.json.gz 'http.favicon.hash:-247388890' 306 shodan parse --fields ip_str,port,org --separator , results.json.gz 307 ``` 308 309 Search filters need a paid or academic key; a free account can look up hosts but not run filtered 310 queries, so budget for that before planning a workflow around it. Installed under Python 3.12 the 311 CLI can fail with `ModuleNotFoundError: pkg_resources` — install `setuptools` into the same 312 environment. Shodan's data is a crawl with a lag of days to weeks, so a result describes what was 313 listening at scan time, and `hostnames` are reverse-DNS, not proof the site is served there. 314 315 ### httpx 316 317 Probes a list of hosts and reports what answered, with the fingerprints you need for linking sites 318 in one pass: favicon hash, JARM, title, technology, body hash. The natural next step after 319 `subfinder`, and the way you generate the favicon hash that Shodan then searches. 320 321 ```bash 322 # install: ProjectDiscovery's httpx, NOT the Python HTTP library of the same name 323 go install -v github.com/projectdiscovery/httpx/cmd/httpx@latest 324 325 # what is actually live, with status and title 326 subfinder -d example.com -silent | httpx -silent -sc -title 327 328 # the fingerprint set, as JSON you can diff between estates 329 httpx -l subs.txt -json -favicon -jarm -tech-detect -title -web-server -o fingerprints.jsonl 330 331 # the favicon hash on its own — this is the number you paste into Shodan 332 httpx -u https://example.com -favicon -silent 333 334 # group hosts by response-body hash to find the ones serving identical pages 335 httpx -l subs.txt -hash mmh3 -silent | sort -k2 | uniq -c -f1 | sort -rn | head 336 337 # whether a CDN is hiding the origin, and whose network it is 338 httpx -l subs.txt -silent -cdn -ip -asn 339 340 # screenshots, for the record rather than the analysis 341 httpx -l subs.txt -screenshot -silent -o shots.txt 342 343 # slow it down: this is the one tool on this page that touches the target 344 httpx -l subs.txt -rate-limit 5 -silent -sc 345 ``` 346 347 This is active: every probe is a request from your address to their server, logged at their end. 348 Everything above it on this page is passive, so do the passive work first and know what you are 349 looking for before you knock. A favicon hash is only a link when the favicon is distinctive — 350 thousands of hosts share the default WordPress and cPanel icons, so check the hash's global count 351 in Shodan before treating a match as meaningful. 352 353 ### Wayback Machine CDX API 354 355 The archive's index, and far more useful than the calendar view: it lists every capture of a URL 356 pattern with timestamp, status code and content digest, which turns "what did this site used to 357 say" into a diffable dataset. 358 359 ```bash 360 CDX='https://web.archive.org/cdx/search/cdx' 361 362 # every distinct capture of a path, collapsed so identical content appears once 363 curl -s -G "$CDX" --data-urlencode 'url=example.com/contact*' \ 364 --data-urlencode 'output=json' --data-urlencode 'collapse=digest' 365 366 # the first and last capture, i.e. when the site appeared and when it stopped 367 curl -s -G "$CDX" --data-urlencode 'url=example.com' --data-urlencode 'output=json' \ 368 --data-urlencode 'limit=1' 369 curl -s -G "$CDX" --data-urlencode 'url=example.com' --data-urlencode 'output=json' \ 370 --data-urlencode 'limit=-1' 371 372 # every URL ever archived under the domain — the archive as a site map 373 curl -s -G "$CDX" --data-urlencode 'url=example.com/*' --data-urlencode 'output=json' \ 374 --data-urlencode 'fl=original' --data-urlencode 'collapse=urlkey' | jq -r '.[1:][] | .[0]' 375 376 # only captures that returned 200, within a date window 377 curl -s -G "$CDX" --data-urlencode 'url=example.com/*' --data-urlencode 'output=json' \ 378 --data-urlencode 'filter=statuscode:200' --data-urlencode 'from=2018' --data-urlencode 'to=2020' 379 380 # pull one archived page as it was served, without the archive's own toolbar 381 curl -s 'https://web.archive.org/web/20190101120000id_/http://example.com/contact' 382 383 # diff two captures of the staff page, which is where removed names live 384 for ts in 20180301000000 20210301000000; do 385 curl -s "https://web.archive.org/web/${ts}id_/http://example.com/team" > "team-$ts.html" 386 done 387 diff team-20180301000000.html team-20210301000000.html | head -40 388 ``` 389 390 Use `https` — the `http` form of the CDX endpoint returns nothing rather than redirecting. The 391 `id_` suffix on a snapshot URL serves the original bytes, which is what you want for forensics and 392 for hashing. The archive honours removal requests and `robots.txt` retroactively in places, so 393 absence is not evidence of absence; [archive.today](https://archive.ph/) catches things Wayback 394 missed and is harder to have taken down. See 395 [archiving and evidence](/sheets/osint/archiving-and-evidence) for making your own captures stand 396 up. 397 398 ### The Information Laundromat 399 400 Web only, free, and purpose-built for the question the rest of this page answers one signal at a 401 time: are these sites run by the same people. Paste in a URL or a block of content and it compares 402 both the text and the technical indicators across its index. 403 404 Use it at [informationlaundromat.com](https://informationlaundromat.com/) — note the hyphenated 405 spelling of the domain no longer resolves. Two modes matter: 406 407 ```text 408 Content similarity paste an article; it finds the other sites republishing the same text, 409 which is how a syndication network becomes visible 410 Metadata similarity give it a domain; it compares analytics IDs, ad IDs, CDN and registrar 411 fingerprints against its index and scores the overlap 412 Read the result as: which indicator matched, not the headline score — a shared Google 413 Analytics ID is near-conclusive, a shared Cloudflare nameserver is noise 414 Capture: the matched indicator, its value, and the date you ran it 415 ``` 416 417 The scoring blends strong and weak indicators, so a high score built entirely on shared hosting 418 means nothing. Open the per-indicator breakdown every time, and confirm the strong ones by hand — 419 a claimed analytics ID should be visible in the page source or in a 420 [PublicWWW](https://publicwww.com/) or [Grep.app](https://grep.app/) search before you rely on it. 421 422 ### Tools that have moved since this area settled 423 424 - **Censys** is retiring Legacy Search (`search.censys.io`) and its v1/v2 APIs through 2026, in 425 favour of the Platform API at `https://api.platform.censys.io/v3/global/`. Scripts written 426 against `/api/v2/hosts/...` need rewriting; certificate observation endpoints have already 427 dropped from realtime to a six-hourly refresh. 428 - **Amass** is at v5 and no longer works the way most cheatsheets describe. Passive is now the 429 default, so `-passive` is a deprecated no-op and `-active` is the flag that adds zone transfers 430 and certificate grabs. Results land in its own asset database, read back with 431 `amass subs` and `amass viz` rather than printed by `amass enum`. For plain passive subdomain 432 lists, `subfinder` is the simpler tool. 433 - **DNSDumpster** now requires an account for anything beyond a handful of lookups a day. The 434 keyless equivalents are crt.sh and `subfinder`. 435 436 ## Infrastructure tooling at a glance 437 438 | Tool | Best use | Network contact | Evidence boundary | 439 | --- | --- | --- | --- | 440 | RDAP / `whois` | Current registration dates, status and delegation | Registry or RDAP proxy | Current record, not historical ownership | 441 | `dig` | Live DNS, delegation and mail-provider discovery | Resolver; `AXFR` contacts the authoritative server | What resolves now | 442 | crt.sh | Certificate names and issuance timeline | Third-party CT index | A certificate was logged, not that the host served traffic | 443 | `subfinder` | Broad passive subdomain collection with source attribution | Third-party passive sources | Historical candidates until resolved | 444 | urlscan.io | Archived browser requests, screenshots and shared identifiers | Third-party archive; a new scan visits the target | One browser run from one location | 445 | Shodan | Favicon, certificate and response pivots across internet scans | Third-party scan index | What Shodan observed at scan time | 446 | `httpx` | Live status, title, hashes, JARM and screenshots | Direct requests to every supplied host | Active observation from your address | 447 | Wayback CDX | Historical URLs, capture dates and content digests | Internet Archive | Archived captures only; gaps are expected | 448 449 ## Tool reference 450 451 | Tool | What it does | Cost | 452 | --- | --- | --- | 453 | [Distill](https://distill.io/) | Distill is a website change monitoring tool that allows users to track changes on web pages. | partly free | 454 | [DNS History](http://completedns.com/) | Collection of historical DNS information. | free | 455 | [Domain Research Suite](https://drs.whoisxmlapi.com/) | Domain Research Suite provides tools to obtain registration and ownership data for domain names, along with historical search and reverse lookup… | paid | 456 | [DomainTools Whois Lookup](https://whois.domaintools.com/) | DomainTools Whois provides detailed domain name registration information, and can be used to investigate details about domains or IP addresses. | partly free | 457 | [Geo Data Tool](https://www.geodatatool.com/) | IP geolocation service to identify the location and other technical information associated to IP addresses. | free | 458 | [Grep.app](https://grep.app/) | grep.app is a free web-based search engine that allows users to search the contents of public GitHub repositories. | free | 459 | [ICANN Lookup](https://lookup.icann.org/) | This tool allows you to search for the current registration data of internet domains. | free | 460 | [IDN Checker](https://holdintegrity.com/checker) | IDN Checker detects visually similar versions of a domain. | free | 461 | [Intelx](http://intelx.io/) | Find user details in data breaches | partly free | 462 | [Moz Link Explorer](http://moz.com/link-explorer) | Analyse the links of any website. | free | 463 | [PublicWWW](https://publicwww.com/) | PublicWWW is a source code search engine that allows you to search for any alphanumeric snippet, signature, or keyword within the HTML, JavaScript, and… | partly free | 464 | [Shodan](https://www.shodan.io/) | A search engine for internet-connected devices, from webcams to databases. | partly free | 465 | [The Information Laundromat](https://informationlaundromat.com) | A tool for analyzing content replication and site architecture to detect information laundering. | free | 466 | [Urlscan](https://urlscan.io/) | urlscan.io is an online tool that allows investigators to analyse, monitor, and document websites in real time. | free | 467 | [Wayback Machine](https://web.archive.org/) | The Wayback Machine is the Internet Archive's free tool for viewing and saving archived web pages, with over a trillion pages captured, widely used for… | free | 468 | [Web Archives](https://github.com/dessant/web-archives) | A browser extension to view archived and cached versions of a website on multiple archiving sites. | partly free | 469 | [What CMS](https://whatcms.org/) | WhatCMS is a web-based tool for anyone needing information about the technologies behind any website, including the content management system (CMS)… | partly free | 470 | [Whoxy](https://www.whoxy.com/) | Whoxy is a domain search engine or "whois lookup" tool to find (the history of) registration information on a domain, such as the registrar, the status of… | partly free | 471 472 ## Pitfalls 473 474 - **Privacy shields hide almost everything current.** History is where the answer is. 475 - **Shared hosting proves nothing.** Thousands of unrelated domains share a CDN IP. Only a 476 dedicated or small shared host is meaningful. 477 - **CDNs hide the origin.** Cloudflare in front of a site means the IP you see is Cloudflare's. 478 - **Active scanning is not passive.** `nmap` against the target leaves logs. Urlscan and CT logs do 479 not. 480 - **Archives have gaps and honour exclusions.** Absence from Wayback is not absence from the web. 481 482 ## Worked example 483 484 One datum: the domain **example-news-daily.com**, which published a story you are checking. The 485 question is who runs it and what else they run. 486 487 1. **Dates first.** `curl -s -L 'https://rdap.org/domain/example-news-daily.com'` gives a 488 registration event in March 2024 and a privacy-shielded registrant. A site presenting itself as 489 a long-running local paper, registered eighteen months ago, is already a finding. 490 2. **Names it has had.** crt.sh for `%.example-news-daily.com` returns six hostnames, including 491 `old.` and `wp.` subdomains and — more usefully — a certificate from 2024 that also covers 492 `example-city-times.com`. One certificate across two brands is a strong link. 493 3. **Broaden the host list.** `subfinder -d example-news-daily.com -all -oJ -cs` adds four names 494 crt.sh missed and records which source found each, so you can cite them individually. 495 4. **See what is live without being seen.** A urlscan archive search for 496 `domain:example-news-daily.com` returns existing public scans, including the third-party requests the page makes. One is a Google 497 Analytics beacon carrying a property ID. 498 5. **Pivot on the ID.** That ID searched in PublicWWW returns eleven other domains embedding it, 499 and the urlscan query `filename:"UA-…"` returns nine of the same. Overlapping lists from 500 two independent indexes is what makes the network claim defensible. 501 6. **Confirm the shared template.** `httpx -l hosts.txt -json -favicon -jarm -title` shows the same 502 favicon hash across the whole set. That hash in `shodan count` returns 14 hosts globally — small 503 enough that the match means something, which a default-WordPress icon would not. 504 7. **Read the history.** The Wayback CDX index for `example-news-daily.com/*` shows the earliest 505 capture in April 2024 and an `/about` page archived in May 2024 naming two editors; the live 506 `/about` names neither. `diff` of the two captures is the record. 507 8. **Check the network, not just the pair.** The Information Laundromat's metadata comparison on the 508 domain scores the same cluster and surfaces two more domains that share the ad ID but not the 509 analytics ID — a wider ring, confirmed on a second indicator. 510 511 What you can assert: a domain registered in March 2024, sharing a certificate, an analytics 512 property, an ad publisher ID and a distinctive favicon with eleven other sites, with two editor 513 names removed from its own about page in 2024 and preserved in the archive. What you cannot: who 514 those people are. Nothing on this page crosses from infrastructure to identity — that pivot runs 515 through [people search](/sheets/osint/people-search) and 516 [email and phone work](/sheets/osint/email-and-phone). 517 518 What would falsify it: a fresh page capture showing that the analytics or ad identifier was 519 injected by a shared third-party template rather than by the operator; certificate history showing 520 that the cross-domain certificate was issued to a hosting platform for unrelated customers; or a 521 global favicon count large enough to make the hash commonplace. The load-bearing measurements are 522 the account-linked analytics and ad identifiers. The shared IP and favicon are corroboration, not 523 identity evidence, and must be discarded if either is common outside the eleven-site cluster. 524 525 ## Broader catalogues 526 527 - [Domain Name OSINT](https://tools.osintnewsletter.com/tool-categories/domain-name-osint) 528 - [Network Infrastructure OSINT](https://tools.osintnewsletter.com/tool-categories/network-infrastructure-osint) 529 - [Cyberthreat Intelligence OSINT](https://tools.osintnewsletter.com/tool-categories/cyberthreat-intelligence-osint) 530 531 532 ## More tools 533 534 Further tools for this area from the OSINT Newsletter Tools Library ([Domain Name OSINT](https://tools.osintnewsletter.com/tool-categories/domain-name-osint), [Network Infrastructure OSINT](https://tools.osintnewsletter.com/tool-categories/network-infrastructure-osint), [Cyberthreat Intelligence OSINT](https://tools.osintnewsletter.com/tool-categories/cyberthreat-intelligence-osint)), excluding those already listed above. 535 536 | Tool | What it does | 537 | --- | --- | 538 | [Analyst Research Tools](https://analystresearchtools.com/) | A browser-based investigative research platform pulling together a collection of free lookup tools to help you gather, organise… | 539 | [ARIN](https://www.arin.net/) | A non-profit, regional internet registry that provides public access to IP address, Autonomous System Number (ASN), and network… | 540 | [Blacklist Alert](https://blacklistalert.org/) | Webtool that checks whether an IP address or domain appears on public DNS-based blacklists (DNSBLs). | 541 | [BotScout](https://botscout.com/search.htm) | A bot-signature lookup service that checks names, email addresses & IP addresses against a database of previously identified bot… | 542 | [Censys](https://censys.com/) | An internet scanning and reconnaissance platform that continuously indexes exposed devices, services, and infrastructure across… | 543 | [Central Ops](https://centralops.net/co/) | A free internet reconnaissance and domain investigation platform that brings together multiple lookup tools in one place. | 544 | [Dark Reading](https://www.darkreading.com/) | A cybersecurity news and intelligence platform providing news, analysis, research, and expert insights on cyber threats… | 545 | [DNS Dumpster](https://dnsdumpster.com/) | A free domain research tool that can discover hosts related to a domain. | 546 | [Domain Digger](https://digger.tools/) | A fast, browser-based OSINT tool for uncovering infrastructure linked to a domain. | 547 | [Domain Dossier](https://centralops.net/co/DomainDossier.aspx) | Investigate domain names, IP addresses, and DNS records to understand how a website is set up and who may be behind it. | 548 | [Dorky](https://dork.bugbountyhunting.com/) | A focused Google dorking and search query builder designed to help quickly generate advanced search queries for uncovering… | 549 | [Favicon Hash Generator](https://favicon-hash.kmsec.uk/) | A lightweight tool for generating a Shodan-compatible favicon hash from a website favicon or local favicon file. | 550 | [File Phish](https://greylensresearch.github.io/filephish/) | A query builder that allows you to discover exposed documents, sensitive files, and hidden data across the web in seconds. | 551 | [FindTheScam](https://findthescam.net/) | Checks whether a website looks legitimate or risky by reviewing WHOIS age, HTTPS/SSL, DNS, reputation, and scam-warning signals. | 552 | [FOFA](https://en.fofa.info/) | A cyberspace search engine for the internet of things that lets you discover exposed systems, servers, and network services… | 553 | [Have I Been Squatted?](https://haveibeensquatted.com/) | Checks whether a domain has been registered as a typo-squatted or lookalike site. | 554 | [Hippie OSINT Toolkit](https://osint.hippie.cat/) | An OSINT web toolkit that allows you to reverse search a domain, a TikTok post, an image, or username (and more). | 555 | [Host.io](https://host.io/) | A domain intelligence and infrastructure discovery tool to uncover relationships between websites, IP addresses, hosting… | 556 | [Hudson Rock](https://www.hudsonrock.com/) | A cybercrime intelligence platform that lets you search for compromised credentials, infected machines, and exposed corporate… | 557 | [IBM X-Force Exchange](https://exchange.xforce.ibmcloud.com/) | Threat intelligence platform providing indicators of compromise (IOCs), malware analysis, IP and domain reputation, vulnerability… | 558 | [IntelligenceX](https://intelx.io/) | An OSINT search engine that indexes breached data, leaks, and historical internet records to help uncover hidden links and… | 559 | [IPinfo](https://ipinfo.io/) | A fast, reliable tool for looking up IP addresses and getting key details like location, ISP, and network ownership. | 560 | [JSON Crack](https://jsoncrack.com/) | Visualise JSON data as interactive graphs to quickly understand structure and relationships. | 561 | [Lookyloo](https://lookyloo.circl.lu/capture) | A web forensics tool that captures a webpage and maps every domain, resource, redirect, and third-party connection involved in… | 562 | [Malpedia](https://malpedia.caad.fkie.fraunhofer.de/) | A curated malware intelligence platform used to identify malware families, research samples, actors and related technical… | 563 | [Meawfy](https://meawfy.com/) | A web-based OSINT crawler that searches and indexes publicly accessible files hosted on MEGA.nz using automated discovery… | 564 | [Netlas](https://netlas.io/) | Internet intelligence and attack surface discovery platform to search and analyse publicly exposed infrastructure, domains… | 565 | [OSINT Industries](https://app.osint.industries/) | An all-encompassing OSINT platform that gathers and correlates publicly available digital data such as emails, domains, phone… | 566 | [PhishTank](https://www.phishtank.com/) | A community-driven database of verified phishing websites, helping you check URLs and track phishing campaigns. | 567 | [Pulsedive](https://pulsedive.com/) | A threat-intelligence and indicator-enrichment platform that allows users to investigate domains, IP addresses, and URL. | 568 | [ScamDB](https://www.scamdb.net/) | Community-driven scam intelligence platform used to search and report suspicious websites, phone numbers, email addresses &… | 569 | [ScanMalware](https://scanmalware.com/) | An online URL and website security-analysis platform for investigating potentially malicious or suspicious websites. | 570 | [SynapsInt](https://synapsint.com/) | A web-based search platform that aggregates publicly available data about people, organisations, domains, IP addresses, email… | 571 | [Threat Actor Username Search](https://threatactorusernames.com/) | A simple OSINT tool that checks whether a username has been observed on known cybercriminal forums and underground platforms. | 572 | [Tiny Scan](https://www.tiny-scan.com/) | Scans a website, pulling valuable information including SSL certificates, records, web technologies and HTTP headers in an… | 573 | [VirusTotal](https://www.virustotal.com/gui/home/upload) | An online threat intelligence platform that aggregates over 70 antivirus scanners and blocklisting services to analyse suspicious… | 574 | [WHOIS API](https://whois.whoisxmlapi.com/) | A domain registration tool showing who owns a domain, when it was created, and how it’s configured. | 575 | [Whoisology](https://whoisology.com/) | WHOIS intelligence platform that provides historical domain ownership records, reverse WHOIS searches and domain registration… | 576 | [XResolver](https://xresolver.com/) | A gamer-tag and IP intelligence lookup tool linking Xbox and PlayStation usernames to IP addresses collected from public and… | 577 578 ## Sources 579 580 Both catalogues below are maintained by other people and are considerably larger than 581 this page. Use them as the canonical index; this sheet is a working route through them. 582 583 - [Bellingcat's Online Investigation Toolkit](https://bellingcat.gitbook.io/toolkit) — ~340 tools, each with its own 584 review page covering cost, difficulty, requirements and limitations. 585 - [OSINT Newsletter Tools Library](https://tools.osintnewsletter.com) — ~280 tools, organised by investigative goal. 586 587 Neither publishes a licence, so nothing here is copied from them: tool names, one-line 588 descriptions, cost flags and links are catalogue facts, and the method and commentary are 589 this site's own. See [credits](/credits).