google-dorking.md (36739B)
1 --- 2 title: "Google Dorking" 3 description: "Search-engine operators as a recon primitive: the full operator reference, the gotchas that silently break dorks, recipes by objective, cross-engine translation, OPSEC, and defending your own estate against them." 4 category: enumeration 5 tags: [enumeration, osint, recon, google-dorking, dorks, passive, search-operators, ghdb] 6 tools: [Google, Bing, DuckDuckGo, Yandex, GitHub, Shodan, dorkforge] 7 difficulty: beginner 8 updated: "2026-09-26" 9 --- 10 11 # Google Dorking 12 13 Dorking is using a search engine's own query language to find things its index already holds but nobody meant to publish: config files, backups, admin panels, stack traces, cloud buckets, private keys. No packet reaches the target — you are querying Google's copy of the web, which is why it opens [Stage 00 — Passive External Recon](/sheets/pentest-workflow/passive-external-recon) and why it survives every firewall, WAF and rate limit the target owns. 14 15 The technique is old and the operators are few. What separates a useful dork from noise is knowing which operators still work, how they combine, and where each one silently fails. 16 17 > [!danger] Indexed is not authorised 18 > Every dork on this page returns *links*. Following one to retrieve a config file, enumerate a bucket, or log into a panel is **access**, and access needs written permission — a public URL is not consent, and "it was in Google" has never been a defence. Personal data you collect is regulated (GDPR and equivalents) whether or not the engagement covers it. Use this against your own estate, a scope you hold in writing, or a lab you are entitled to. 19 20 --- 21 22 ## 1. Why it works — you are searching the index, not the site 23 24 A crawler finds a URL, fetches it, and the indexer stores the text, title, URL and file type. Everything dorking does is filter that stored record. Three consequences matter: 25 26 | Fact | Consequence for recon | 27 |---|---| 28 | `robots.txt` is a **crawl** directive, not an access control | A `Disallow:` line stops polite crawlers fetching a path — it does not stop the path being served, and it publishes a list of the paths the owner considers sensitive. Fetch `$DOMAIN/robots.txt` first; it is a free sitemap of the interesting bits. | 29 | Only `noindex` removes a page from results | `X-Robots-Tag: noindex` or a `<meta name="robots" content="noindex">` tag is the actual removal mechanism, and it only works if the crawler is *allowed in* to see it. A `Disallow`ed page that is linked from elsewhere can still appear. | 30 | The index outlives the file | Deleting a file does not delete the record. Titles and snippets persist for weeks; archives persist indefinitely. A dork that returns a 404 still tells you the file existed and what it was called. | 31 | Anything linked once is crawlable | A "secret" URL pasted into a public ticket, a README, a Slack export or a CDN log gets crawled. This is how private buckets and staging hosts end up indexed. | 32 33 The practical reading: **the index is a historical record of what the target served, not a snapshot of what it serves now.** Treat every hit as a lead with a timestamp, not a live finding. 34 35 --- 36 37 ## 2. Anatomy of a dork 38 39 ```text 40 site:example.com ext:sql intext:"CREATE TABLE" -site:docs.example.com 41 └──── scope ───┘ └─ type ┘ └───── content ────┘ └────── exclusion ─────┘ 42 ``` 43 44 Four moving parts, and only the first is mandatory in practice: 45 46 - **Scope** — which hosts count (`site:`). 47 - **Type** — what kind of document (`ext:` / `filetype:`). 48 - **Content** — a literal that only the interesting version of the file contains (`intext:`, `intitle:`, `inurl:`). 49 - **Exclusion** — what to subtract once the first pass is noisy (`-site:`, `-inurl:`). 50 51 Terms and operators are **AND**-ed by default. There is no `AND` keyword — writing one searches for the word "and". 52 53 > [!tip] Build them in that order 54 > Scope first, then type, then narrow with content, then subtract noise. Starting from a long stacked dork you copied off a list gives you zero results and no idea which clause killed it. Start broad, add one operator at a time, and watch the result count move. 55 56 --- 57 58 ## 3. Operator reference 59 60 ### Scope 61 62 | Operator | Does | Example | 63 |---|---|---| 64 | `site:` | Restrict to a host. Matches the host as a **suffix**, so it covers subdomains. | `site:example.com` | 65 | `site:*.` | Force the subdomain tree. Pair with a negative to drop the ones you know. | `site:*.example.com -site:www.example.com` | 66 | `-site:` | Subtract a host. The main tool for forcing new subdomains to the surface. | `-site:blog.example.com` | 67 | `related:` | Sites Google considers similar. **Officially unsupported** — usually returns nothing. | `related:example.com` | 68 69 ### Content position 70 71 | Operator | Matches in | Notes | 72 |---|---|---| 73 | `intitle:` | The `<title>` element | Highest-signal field for fingerprinting an app by its shipped default title. | 74 | `inurl:` | The whole URL, **including the query string** | This is why it finds parameter names: `inurl:redirect=`. Substring match. | 75 | `intext:` | The page body | Forces a term into the body instead of letting it match the title or an inbound link. | 76 | `inanchor:` | Inbound link text | Thin and unreliable now that link signals are de-emphasised. | 77 | `allintitle:` `allinurl:` `allintext:` | Same fields, all following terms | **Do not use.** See the gotchas below. | 78 79 ### File type 80 81 | Operator | Does | Notes | 82 |---|---|---| 83 | `filetype:` | One file extension | Matches the extension Google recorded, not the `Content-Type` header. | 84 | `ext:` | Identical to `filetype:` | Shorter, so stacked dorks stay readable. Both work on Google, Bing and Brave. | 85 86 One extension per operator — stack them with `OR` to cover a family: 87 88 ```text 89 site:$DOMAIN ext:sql OR ext:bak OR ext:old OR ext:backup 90 ``` 91 92 Google indexes a fixed set of document types (`pdf`, `doc`/`docx`, `xls`/`xlsx`, `ppt`/`pptx`, `rtf`, `ps`, `kml`, `txt` and a handful more). A `ext:env` dork works **only because such files were served as plain text and indexed as text** — which is precisely what makes a hit worth reporting. 93 94 ### Logic and grouping 95 96 | Syntax | Does | Trap | 97 |---|---|---| 98 | `"phrase"` | Exact match, in order, no stemming or synonyms | Any literal artefact — a banner, an error string, a key header — belongs in quotes. | 99 | `OR` or `\|` | Disjunction | **Must be uppercase.** Lowercase `or` is a stop word and is silently ignored. | 100 | `-term` | Exclude | No space after the hyphen. Works on operators too: `-site:`, `-inurl:`. | 101 | `( )` | Grouping | Required whenever you mix `OR` with AND-ed terms, or precedence bites you. | 102 | `*` | Whole-token wildcard | Only meaningful **inside quotes**. It is not a substring wildcard — `adm*n` does not work. | 103 | `AROUND(n)` | The two terms within *n* words | Uppercase, no space before the bracket. Tighter than AND, looser than a phrase. | 104 105 ```text 106 site:$DOMAIN (ext:sql OR ext:bak) intext:password # grouped correctly 107 site:$DOMAIN ext:sql OR ext:bak intext:password # ambiguous — do not 108 password AROUND(3) database site:$DOMAIN # proximity 109 "confidential * report" site:$DOMAIN # token wildcard 110 ``` 111 112 ### Time and number 113 114 | Syntax | Does | Caveat | 115 |---|---|---| 116 | `before:YYYY-MM-DD` | Documents Google dates before that day | Filters on Google's *estimate* of the document date, which is frequently wrong for pages without date markup. | 117 | `after:YYYY-MM-DD` | Documents Google dates after that day | Same caveat. Good for narrowing; never good enough to prove publication date. | 118 | `1000..2000` | Any number in the range | Two dots. Legal inside other operators: `inurl:id=1..100`. | 119 120 ### Retired — and what replaced them 121 122 > [!warning] These appear in every dork list on the internet and none of them work 123 > | Operator | Status | Use instead | 124 > |---|---|---| 125 > | `cache:` | Removed by Google in **September 2024**, along with the cached-page links | [web.archive.org](https://web.archive.org), [archive.today](https://archive.today) | 126 > | `link:` | Deprecated **2017**; now a silent no-op that degrades to a keyword search | Search Console (for sites you own) or a commercial backlink tool | 127 > | `info:` | Retired; folded into ordinary search | Just search the bare URL | 128 > | `+term` | Removed **2011** | Quote the term instead: `"term"` | 129 > | `daterange:` | Julian-date legacy syntax, unreliable | `before:` / `after:` | 130 > 131 > A retired operator does not error. Google drops it and searches the remaining words, so your dork returns plausible-looking garbage. This is the single most common reason a copied dork "stops working". 132 133 --- 134 135 ## 4. The gotchas that actually break dorks 136 137 > [!example] Read this section twice — it is most of the skill 138 > | Trap | What happens | Fix | 139 > |---|---|---| 140 > | Lowercase `or` | Treated as a stop word, dropped. Your disjunction became an AND. | Always uppercase `OR`. | 141 > | `allintitle:` / `allinurl:` / `allintext:` | They consume **the rest of the query** — every operator after them is ignored, including `site:`. Your scoped dork quietly went global. | Stack the singular form instead: `intitle:a intitle:b`. | 142 > | Space after the colon | `site: example.com` is parsed as a bare `site:` plus the keyword `example.com`. | Never put a space after an operator's colon. | 143 > | Unquoted multi-word value | `intitle:index of` means `intitle:index` AND the word `of`. | Quote it: `intitle:"index of"`. | 144 > | `*` outside quotes | Ignored entirely. | Only use it inside a quoted phrase. | 145 > | `site:` is a suffix match | `site:example.com` also returns `notexample.com`-style matches on some engines and always includes every subdomain. | Add `-site:` exclusions, or use Yandex's stricter `host:`. | 146 > | Very long stacked dorks | Google historically truncated long queries, and still degrades them — the tail of a 15-operator dork may never be evaluated. | Split into two or three narrower dorks. | 147 > | Personalised results | Signed in, Google reorders and filters results for *you*. Two operators can look broken because your profile suppressed the hits. | Search signed out, in a clean profile. | 148 > | Result count is a lie | The "About N results" figure is an estimate and swings wildly. | Page to the end. The real count is where results stop. | 149 > | Verbatim mode | Google rewrites and expands queries by default. | Append `&tbs=li:1` to the results URL, or use Tools → All results → Verbatim. | 150 151 > [!tip] Two switches worth memorising 152 > `&num=100` on the results URL returns 100 per page instead of 10. `&filter=0` disables the "omitted similar results" collapse, which routinely hides the interesting duplicates on a large site. 153 154 --- 155 156 ## 5. Recipes by objective 157 158 Every dork below is written in the Google dialect with `$DOMAIN`, `$ORG` and `$KEYWORD` as placeholders. Section 6 covers translating them. 159 160 ### Exposed files and configs 161 162 The highest-yield category in the whole practice. 163 164 | Dork | Finds | 165 |---|---| 166 | `site:$DOMAIN ext:env OR ext:cfg OR ext:conf OR ext:ini` | Application config. A served `.env` is plaintext credentials: database URI, mail password, cloud keys, app secret. | 167 | `site:$DOMAIN ext:bak OR ext:old OR ext:backup OR ext:swp` | Editor and deploy leftovers. `index.php.bak` is served as text instead of executed — that is the source code. | 168 | `site:$DOMAIN ext:log OR ext:txt inurl:log` | Logs: internal paths, usernames, session tokens, stack traces. | 169 | `site:$DOMAIN ext:yml OR ext:yaml OR ext:toml inurl:config` | Compose files, CI definitions, `appsettings`, secrets baked into a values file. | 170 | `site:$DOMAIN ext:pem OR ext:key OR ext:ppk OR ext:p12 OR ext:pfx` | Private keys and cert bundles. Any hit is a critical finding on its own. | 171 | `site:$DOMAIN inurl:.git OR inurl:.svn OR inurl:.hg` | Exposed VCS metadata — a readable `.git/` means the whole repo and its history. | 172 | `site:$DOMAIN inurl:wp-config OR inurl:configuration.php OR inurl:settings.py` | Framework config by canonical filename. | 173 174 ### Login and admin panels 175 176 | Dork | Finds | 177 |---|---| 178 | `site:$DOMAIN inurl:admin OR inurl:administrator OR inurl:adminpanel` | The obvious paths. Still hits constantly. | 179 | `site:$DOMAIN intitle:"login" OR intitle:"sign in"` | Title-based, so it catches panels on paths you would never guess. | 180 | `site:$DOMAIN inurl:phpmyadmin OR inurl:adminer OR inurl:pma` | Exposed database consoles, frequently on vendor defaults. | 181 | `site:$DOMAIN intitle:"Dashboard" (inurl:jenkins OR inurl:grafana OR inurl:kibana)` | CI and observability — build secrets, internal hostnames, production telemetry. | 182 | `site:$DOMAIN inurl:/manager/html OR inurl:/jmx-console OR inurl:/axis2` | Java app-server managers. Tomcat manager with default creds is a WAR deploy from RCE. | 183 | `site:$DOMAIN inurl:owa OR inurl:/rdweb OR inurl:citrix OR inurl:vpn` | Remote-access front doors — the surface spraying actually targets. | 184 185 ### Open directory listings 186 187 | Dork | Finds | 188 |---|---| 189 | `site:$DOMAIN intitle:"index of"` | The canonical dork. Apache, nginx and IIS all emit a title starting this way. | 190 | `site:$DOMAIN intitle:"index of" (backup OR bak OR old OR archive)` | Listings that name themselves as a dumping ground. | 191 | `site:$DOMAIN intitle:"index of" "parent directory" (.sql OR .zip OR .tar.gz)` | Listings with a dump or archive sitting in them. | 192 | `site:$DOMAIN "Directory Listing For" -inurl:html` | Tomcat and others whose listing wording differs from Apache's. | 193 194 ### Cloud storage 195 196 | Dork | Finds | 197 |---|---| 198 | `site:s3.amazonaws.com "$ORG"` | S3 buckets whose listing mentions the organisation. | 199 | `site:$DOMAIN inurl:s3.amazonaws.com` | Pages on the target that link into S3 — the fastest way to learn the bucket naming convention. | 200 | `inurl:blob.core.windows.net "$ORG"` | Azure Blob containers. | 201 | `inurl:storage.googleapis.com "$ORG"` | Google Cloud Storage objects. | 202 | `site:$DOMAIN inurl:r2.dev OR inurl:digitaloceanspaces.com` | Smaller providers — less scrutiny, misconfigured more often. | 203 | `"$ORG" (site:trello.com OR site:notion.site)` | Public boards and docs. Trello leaks credentials and architecture diagrams at a remarkable rate. | 204 205 ### Credentials, keys and tokens 206 207 | Dork | Finds | 208 |---|---| 209 | `site:$DOMAIN intext:"BEGIN RSA PRIVATE KEY"` | Key headers pasted into a wiki, a ticket, a README. | 210 | `site:$DOMAIN intext:"AKIA"` | AWS access key IDs all begin `AKIA` — a distinctive literal that rarely false-positives. | 211 | `"$ORG" ("xoxb-" OR "ghp_" OR "sk_live_")` | Slack, GitHub and Stripe token prefixes. Structurally unique, so a hit is almost never noise. | 212 | `site:$DOMAIN (ext:env OR ext:yml) intext:PASSWORD` | Config files with a password assignment. | 213 | `site:pastebin.com OR site:justpaste.it "$DOMAIN"` | Paste sites — where dumps, configs and credential lists get parked. | 214 215 > [!danger] A live key is a finding, not a login 216 > Do not authenticate with anything you find. Record the URL, report it immediately so it can be rotated, and treat your own notes as sensitive material for the rest of the engagement. See [Credential Hunting](/sheets/enumeration/credential-hunting), [Gitleaks](/sheets/enumeration/2-4-cheatsheet-gitleaks) and [TruffleHog](/sheets/enumeration/2-5-cheatsheet-trufflehog) for the systematic version of this. 217 218 ### Documents and metadata 219 220 | Dork | Finds | 221 |---|---| 222 | `site:$DOMAIN ext:pdf OR ext:docx OR ext:xlsx OR ext:pptx` | The corpus. Pull it, then run `exiftool` over the lot. | 223 | `site:$DOMAIN ext:pdf (confidential OR internal OR "not for distribution")` | Documents labelled restricted that are nonetheless public. | 224 | `site:$DOMAIN ext:xlsx (employee OR salary OR roster)` | Spreadsheets — the format people paste staff lists into. | 225 | `site:$DOMAIN ext:vsd OR ext:vsdx OR ext:drawio` | Network and architecture diagrams. | 226 227 ```bash 228 # the half of this that isn't the dork: metadata gives you usernames and internal paths 229 exiftool -Author -Creator -Producer -Company *.pdf | sort -u 230 exiftool -a -G1 -s report.docx | grep -Ei 'author|company|template|lastmodifiedby' 231 ``` 232 233 ### Error messages and stack traces 234 235 | Dork | Finds | 236 |---|---| 237 | `site:$DOMAIN intext:"You have an error in your SQL syntax"` | The literal MySQL parser message — confirms injectable input. | 238 | `site:$DOMAIN intext:"Traceback (most recent call last)"` | Python tracebacks. Django debug pages additionally dump settings and environment. | 239 | `site:$DOMAIN intitle:"Whoops, looks like something went wrong"` | Laravel's debug page, which shows environment variables verbatim. | 240 | `site:$DOMAIN intext:"Fatal error" OR intext:"Uncaught exception"` | PHP and Java fatals — usually carry the absolute filesystem path. | 241 | `site:$DOMAIN intext:"Server Error in" intext:"Stack Trace"` | ASP.NET yellow-screen-of-death with tracing left on. | 242 | `site:$DOMAIN intitle:"phpinfo()"` | Full build config, loaded modules, absolute paths, sometimes environment secrets. | 243 244 ### Databases, backups and dumps 245 246 | Dork | Finds | 247 |---|---| 248 | `site:$DOMAIN ext:sql intext:"INSERT INTO"` | SQL dumps confirmed by their own statement syntax, not just an extension. | 249 | `site:$DOMAIN ext:sql intext:"CREATE TABLE" intext:password` | Dumps that include a credentials table. | 250 | `site:$DOMAIN (ext:zip OR ext:tar OR ext:gz) (backup OR dump OR export)` | Archived backups in the webroot. | 251 | `site:$DOMAIN inurl:backup OR inurl:dump OR inurl:export` | Path-based — catches the directory even when the file was not indexed. | 252 253 ### Injectable parameters 254 255 A target list for later, authorised testing. Nothing here is a finding by itself. 256 257 | Dork | Vulnerability class | 258 |---|---| 259 | `site:$DOMAIN inurl:id= OR inurl:pid= OR inurl:cat=` | SQL injection, IDOR | 260 | `site:$DOMAIN inurl:file= OR inurl:page= OR inurl:path=` | LFI, path traversal — see [LFI](/sheets/enumeration/lfi) | 261 | `site:$DOMAIN inurl:url= OR inurl:redirect= OR inurl:next=` | Open redirect, the front half of many SSRF chains | 262 | `site:$DOMAIN inurl:cmd= OR inurl:exec= OR inurl:query=` | Command injection, reflected XSS | 263 | `site:$DOMAIN inurl:debug=true OR inurl:admin=1` | Flags that flip the app into a verbose or privileged mode | 264 265 ### People, emails and usernames 266 267 | Dork | Finds | 268 |---|---| 269 | `site:linkedin.com/in "$ORG"` | Staff profiles — names plus role, which is what derives the username scheme. | 270 | `site:$DOMAIN intext:"@$DOMAIN"` | The org's own pages leaking its address format (`first.last@`, `flast@`). | 271 | `"$ORG" (site:github.com OR site:gitlab.com) intext:"@$DOMAIN"` | Developers using a work address in commits or profiles. | 272 | `"$ORG" (site:stackoverflow.com OR site:serverfault.com)` | Engineers asking about their own stack, pasting real config with real hostnames. | 273 | `"$ORG" ("we use" OR "powered by") -site:$DOMAIN` | Tech stack from job ads, case studies, conference talks. | 274 275 > [!warning] Personal data is regulated independently of your scope 276 > Collect the minimum that supports the finding, store it with the engagement material, and delete it when the engagement closes. 277 278 ### Subdomains and forgotten hosts 279 280 | Dork | Finds | 281 |---|---| 282 | `site:*.$DOMAIN -site:www.$DOMAIN` | The core subdomain dork. | 283 | `site:*.$DOMAIN (dev OR staging OR test OR uat OR qa)` | Non-production — weaker credentials, debug on, real data. | 284 | `site:*.$DOMAIN (vpn OR remote OR portal OR intranet)` | Access infrastructure and internal-facing hosts. | 285 | `"$DOMAIN" -site:$DOMAIN` | Third-party references: partners, status pages, monitoring, archived copies. | 286 287 Dorking is the *weakest* of the subdomain sources — use it to catch what certificate transparency and DNS missed, not as the primary pass. See [Amass](/sheets/enumeration/amass) and [Stage 00](/sheets/pentest-workflow/passive-external-recon) for the real enumeration. 288 289 ### Archived and removed pages 290 291 | Dork | Finds | 292 |---|---| 293 | `site:web.archive.org "$DOMAIN"` | Wayback captures Google indexed. | 294 | `site:archive.today "$DOMAIN"` | The other archive, which captures pages Wayback refuses. | 295 | `site:$DOMAIN before:2022-01-01 "$KEYWORD"` | Older material still on the live site. | 296 297 ```bash 298 # the archive's own API beats dorking it — every captured URL, deduplicated 299 curl -s "http://web.archive.org/cdx/search/cdx?url=$DOMAIN*&output=text&fl=original&collapse=urlkey" \ 300 | sort -u > wayback-urls.txt 301 # pull the ones that look like config or secrets out of the pile 302 grep -Ei '\.(env|sql|bak|old|log|yml|ini|conf)($|\?)' wayback-urls.txt 303 ``` 304 305 --- 306 307 ## 6. Cross-engine translation 308 309 Google is the reference dialect, not the only one. Yandex and Bing routinely index hosts Google never crawled, so a subdomain dork is worth running through at least two engines. 310 311 | Intent | Google | Bing | DuckDuckGo | Yandex | Brave | 312 |---|---|---|---|---|---| 313 | Scope to host | `site:` | `site:` | `site:` | `site:` / `host:` / `rhost:` | `site:` | 314 | File type | `filetype:` `ext:` | `filetype:` `ext:` | `filetype:` | `mime:` | `filetype:` | 315 | In title | `intitle:` | `intitle:` | `intitle:` | `title:` | `intitle:` | 316 | In URL | `inurl:` | `inurl:` | `inurl:` | `url:` | `inurl:` | 317 | In body | `intext:` | `inbody:` | — | *(default)* | — | 318 | Exact phrase | `"..."` | `"..."` | `"..."` | `"..."` | `"..."` | 319 | Exclude | `-term` | `-term` | `-term` | `-term` | `-term` | 320 | Either/or | `OR` | `OR` | `OR` | `OR` | `OR` | 321 | Proximity | `AROUND(n)` | — | — | `/+n` | — | 322 | Date bound | `before:` `after:` | *(UI filter)* | *(UI filter)* | `date:` | *(UI filter)* | 323 324 **Engine-native operators with no Google equivalent — worth the detour on their own:** 325 326 | Engine | Operator | Why you would bother | 327 |---|---|---| 328 | Bing | `ip:203.0.113.10` | Every site Bing indexed on that address — a free reverse-IP lookup, and often the rest of a shared estate behind one origin. | 329 | Bing | `contains:sql` | Pages that **link to** a `.sql` file rather than the file itself — catches the index page of a backup directory that was never indexed. | 330 | Bing | `url:example.com/path` | Tests whether one exact URL is in the index. | 331 | Yandex | `rhost:com.example.*` | Reversed-domain match across the whole tree — stricter and more predictable than Google's suffix matching. | 332 | Yandex | `mime:pdf` | Yandex's `filetype:`, keyed on recorded MIME type. | 333 | DuckDuckGo | `!bangs` (`!gh`, `!so`) | Not operators — redirects to another site's own search. Handy, but the query leaves DDG. | 334 335 > [!note] Brave and DuckDuckGo are the low-friction options 336 > Neither requires an account and neither personalises results, which makes them the right place to start from a clean browser. Their operator support is genuinely narrower than Google's and less documented — when a dork returns nothing there, re-run it on Google before concluding the target is clean. 337 338 --- 339 340 ## 7. Other dialects worth knowing 341 342 ### GitHub code search 343 344 A different language entirely, and the single richest source of leaked secrets. Requires being signed in — use a throwaway account. 345 346 | Operator | Example | Notes | 347 |---|---|---| 348 | `org:` | `org:$ORG path:.env` | Scope to an organisation's public repos. | 349 | `repo:` | `repo:owner/name "password"` | One repository. | 350 | `path:` | `path:.github/workflows` | Replaced the old `filename:`. Globs: `path:**/config/*.yml`. | 351 | `language:` | `"$DOMAIN" language:yaml` | Narrows a noisy content search. | 352 | `symbol:` | `symbol:connectDatabase` | Definitions rather than mentions. | 353 354 ```text 355 org:$ORG path:.env 356 org:$ORG "BEGIN RSA PRIVATE KEY" 357 org:$ORG path:.github/workflows 358 "$DOMAIN" ("192.168." OR "10.0." OR ".local") 359 "$DOMAIN" "password" language:yaml 360 ``` 361 362 The last two are the interesting ones: they find *other people's* repos that leak the target's internal addressing and credentials — contractors, ex-employees, integration partners. 363 364 ### Shodan and Censys 365 366 These index service banners, not web pages, so the web operators do not apply at all. Covered properly in [Shodan](/sheets/enumeration/shodan); the translation for a dorking reflex: 367 368 | You want | Shodan | 369 |---|---| 370 | Hosts presenting the org's certificate | `ssl.cert.subject.cn:$DOMAIN` | 371 | Pages referencing the domain | `http.html:"$DOMAIN"` | 372 | A specific page title | `http.title:"index of"` | 373 | The org's netblocks | `org:"$ORG"` or `net:203.0.113.0/24` | 374 375 The origin-IP hunt behind a CDN is `ssl.cert.subject.cn:` — the origin still presents the real certificate. 376 377 --- 378 379 ## 8. The Google Hacking Database 380 381 The [GHDB](https://www.exploit-db.com/google-hacking-database) is Exploit-DB's curated archive of dorks, started by Johnny Long, now thousands of entries across categories: files containing passwords, sensitive directories, vulnerable servers, error messages, footholds. 382 383 How to use it without wasting an afternoon: 384 385 - **Filter by date.** Most of the archive is a decade old and targets software nobody runs. Sort newest first. 386 - **Read it as a pattern library, not a copy-paste list.** The value of `intitle:"index of" "service.pwd"` is not that exact string — it is that vendor-default filenames are the way in. Adapt the shape to the stack you actually found. 387 - **Scope everything.** Almost every GHDB entry is unscoped and returns strangers' infrastructure. Add `site:$DOMAIN` before you run it, every time. 388 389 --- 390 391 ## 9. OPSEC 392 393 > [!danger] You are the one being logged 394 > Dorking sends nothing to the target — but it sends everything to the search engine, tied to your account, your IP and your browser fingerprint. The queries themselves describe your intent with unusual clarity. 395 396 | Rule | Why | 397 |---|---| 398 | Never dork signed in to a real account | The query history is retained, attributable, and personalises your results into uselessness. Use a clean profile or a private window. | 399 | Use a throwaway for GitHub code search | It is the one dialect that mandates an account. Do not attach your identity to a secrets hunt. | 400 | Expect CAPTCHAs, and slow down | A burst of operator-heavy queries from one IP trips Google's automation detection. Once you are CAPTCHA'd, that IP is degraded for hours. | 401 | Do not click through to live secrets | Fetching the `.env` is a request to the target — logged, attributable, and outside "passive" recon. The snippet in the result usually contains enough to report. | 402 | Treat your notes as sensitive | A session log full of working dorks and hit URLs is a map of the target's exposure. Store it with the engagement, encrypt it, delete it on close. | 403 | Report live keys immediately | Rotation is time-sensitive. Do not sit on a working credential until the report. | 404 405 --- 406 407 ## 10. Automation, rate limits and the ToS question 408 409 Scripted scraping of Google results violates its Terms of Service and fails in practice: the automation detection is good, the HTML changes, and the CAPTCHA is unsolvable at scale without paying someone to solve it. Every "google dork scanner" on GitHub has the same lifecycle — works for a week, then returns empty pages forever. 410 411 The options that actually work: 412 413 | Approach | Reality | 414 |---|---| 415 | **Do it by hand** | Slow, unblockable, and you read each result properly. For a single engagement this is genuinely the right answer. | 416 | **Programmable Search Engine JSON API** | Google's sanctioned path. Free tier is ~100 queries/day, then paid. Supports most operators. Results are scoped to a custom engine you configure, which can be set to search the whole web. | 417 | **A commercial SERP API** | Pays someone else to absorb the blocking. Fine for volume; check that sending client data to a third party is acceptable under your engagement's confidentiality terms. | 418 | **Shodan / Censys APIs** | Purpose-built, properly documented, generous limits. If the thing you want is a service rather than a document, stop dorking and query these. | 419 | **Bing Search APIs** | Microsoft has been winding these down — verify current availability before building anything on them. | 420 421 A tool that *builds* dorks and hands them to your browser has none of these problems, because you are still the one searching. That is what [dorkforge](#13-dorkforge--the-companion-script) does. 422 423 --- 424 425 ## 11. Blue team — defending your own estate 426 427 Dorking is trivially turned around: run it against yourself, on a schedule. 428 429 | Control | What it actually does | 430 |---|---| 431 | `X-Robots-Tag: noindex` header, or `<meta name="robots" content="noindex">` | **The** removal mechanism. Note the trap: a path blocked in `robots.txt` cannot be crawled, so the crawler never sees the `noindex` — the page can stay indexed forever. Allow the crawl, serve the `noindex`, then block it once it has dropped out. | 432 | Search Console removal tool | Fast takedown for a site you own (~6 months), which buys time while you fix the underlying exposure. It is a suppression, not a fix. | 433 | `robots.txt` | A crawl hint, and a published list of your sensitive paths. Never treat it as access control; assume attackers read it first. | 434 | Disable autoindex | `Options -Indexes` (Apache), `autoindex off` (nginx). Kills the entire `intitle:"index of"` class. | 435 | Block dotfiles and backup extensions at the edge | Return 404 for `\.(env\|git\|bak\|old\|sql\|swp)$` before the request reaches the app. | 436 | Turn off debug in production | `APP_DEBUG=false`, `customErrors="On"`, `DEBUG = False`. Removes the entire error-message class in one change. | 437 | Bucket ACLs and public-access blocks | Cloud storage exposure is a configuration problem; no search-engine control fixes it. | 438 | Continuous secret scanning | Gitleaks/TruffleHog in CI plus GitHub push protection catches the commit before it is ever indexed. | 439 | Dork yourself on a schedule | The blue-team use of this whole page. Automate the handful of dorks that matter for your estate and alert on new hits. | 440 441 ```bash 442 # a minimal self-audit, run monthly against your own domain 443 for d in 'ext:env' 'ext:sql' 'ext:bak' 'intitle:"index of"' 'intext:"BEGIN RSA PRIVATE KEY"'; do 444 printf '%s\n' "site:$DOMAIN $d" 445 done 446 # feed them to the Programmable Search JSON API and diff the hit list against last month's 447 ``` 448 449 --- 450 451 ## 12. Practising safely 452 453 - **Your own domains.** The only fully unambiguous target, and the one where findings are actionable. 454 - **`google-gruyere.appspot.com`**, **`testphp.vulnweb.com`**, **`demo.testfire.net`** — deliberately vulnerable hosts published for training. Verify the terms on each before touching it. 455 - **The GHDB, read but not run.** Study the shapes; run them scoped to yourself. 456 - **CTF and lab platforms** with an OSINT category, where the scope is explicit. 457 458 > [!warning] Unscoped dorks return strangers' infrastructure 459 > A dork like `intitle:"webcamXP"` returns hardware belonging to people who never consented to anything. Looking at the result list is one thing; connecting to a device is unauthorised access in essentially every jurisdiction. Scope with `site:` or do not run it. 460 461 --- 462 463 ## 13. dorkforge — the companion script 464 465 **[dorkforge.py](/downloads/enumeration/dorkforge.py)** ([SHA-256](/downloads/enumeration/dorkforge.py.sha256)) is a single-file interactive workbench for everything on this page: it asks what you are hunting for, builds the dork, **explains every operator it used**, translates it for the engine you picked, and hands it to your clipboard, your browser, or an engagement log. 466 467 It ships the full library from section 5 — **14 objectives, 81 recipes, 32 operators** — and it sends no traffic to any target. The searching still happens in your browser, under your account, from your IP, which is why it opens with the authorisation gate and records a scope reference in every session log. 468 469 ### Run it 470 471 It uses [PEP 723](https://peps.python.org/pep-0723/) inline metadata, so `uv` resolves `rich` and `questionary` into a throwaway environment on first run. Nothing to install, nothing to clean up. 472 473 ```bash 474 # download + verify 475 curl -sO https://cheatsheet.daemon-sec.xyz/downloads/enumeration/dorkforge.py 476 curl -sO https://cheatsheet.daemon-sec.xyz/downloads/enumeration/dorkforge.py.sha256 477 shasum -a 256 -c dorkforge.py.sha256 # expect: dorkforge.py: OK 478 479 # run it 480 uv run dorkforge.py 481 ``` 482 483 Expected SHA-256: 484 485 ```text 486 3173df88ab3102ebc4999985fcb881298f3f51374170592e5ac95d0674210193 487 ``` 488 489 Or make it executable and let the shebang do the work — `#!/usr/bin/env -S uv run --script`: 490 491 ```bash 492 chmod +x dorkforge.py 493 ./dorkforge.py 494 # keep it on PATH like the rest of your toolkit 495 install -m 755 dorkforge.py ~/.local/bin/dorkforge 496 ``` 497 498 > [!note] Don't have uv? 499 > `curl -LsSf https://astral.sh/uv/install.sh | sh` (or `brew install uv`). Failing that, `pip install rich questionary` and run it with plain `python3` — the script degrades to that path and tells you so. 500 501 ### The flow 502 503 ```text 504 ▪ DÆMON//SEC 505 dorkforge v1.0.0 search-operator workbench 506 ──────────────────────────────────────────────────────────── 507 508 ? What are you hunting for Exposed files & configs 509 ? Target domain: example.com 510 ? Which engine? Google 511 ? Which dork? site:example.com ext:env OR ext:cfg OR ext:conf OR ext:ini 512 513 ┌─ Google ─────────────────────────────────────────────────┐ 514 │ │ 515 │ site:example.com ext:env OR ext:cfg OR ext:conf OR ext:ini│ 516 │ │ 517 │ Application config. A served .env is credentials in │ 518 │ plaintext - database URI, mail password, cloud keys. │ 519 │ │ 520 └────────────────────────────────────────────────────────────┘ 521 522 OPERATOR WHAT IT DOES GOOGLE 523 ────────────────────────────────────────────────────────────── 524 site: Only return pages whose host matches. ✓ 525 ext: Synonym for filetype: on Google and Bing. ✓ 526 OR Match either side. Must be uppercase. ✓ 527 528 ? Now what? Copy to clipboard / Open in browser / Save to log 529 ``` 530 531 Pick a non-Google engine and it rewrites the dork and tells you what changed: 532 533 ```text 534 $ uv run dorkforge.py --objective listing --domain example.com --engine yandex 535 536 site:example.com title:"index of" 537 dialect notes 538 ▪ intitle: rewritten to title: for Yandex 539 ``` 540 541 ### Usage 542 543 | Command | Does | 544 |---|---| 545 | `uv run dorkforge.py` | Interactive. The main path. | 546 | `dorkforge --session engagement.md` | Append every dork you build to a markdown log, with the URL, the reasoning and the translation notes. | 547 | `dorkforge --objective files --domain example.com` | Non-interactive: print every recipe for one objective. | 548 | `dorkforge --objective code --org acme --engine github` | Same, in GitHub's dialect. | 549 | `dorkforge --list` | The whole library — 14 objectives, 81 recipes — to a pager or a file. | 550 | `dorkforge --operators` | The operator reference table, with live / unreliable / retired status. | 551 | `dorkforge --explain site:` | The long form on any one operator, plus which engines support it. | 552 | `dorkforge --doctor` | Check clipboard helper, browser, colour support and state directory. | 553 | `dorkforge --ascii` | ASCII markers instead of Nerd Font glyphs. | 554 | `dorkforge --reset-ack` | Forget the authorisation acknowledgement and ask again. | 555 556 **Flags:** `--domain` `--org` `--keyword` `--ext` pre-fill the scope prompts. `--engine` picks the dialect (`google` `bing` `ddg` `yandex` `brave` `github` `shodan`). `--no-color` for piping. 557 558 ### The session log 559 560 `--session` writes engagement-ready markdown as you work — the dork, why it was built, the live URL, and anything that did not survive translation: 561 562 ~~~markdown 563 # dorkforge session 564 565 - **Started:** 2026-09-26 14:21 BST 566 - **Authorised scope:** ACME-2026-114 567 - **Tool:** dorkforge 1.0.0 568 569 ## 1. Google 570 571 Application config. A served .env is credentials in plaintext. 572 573 ```text 574 site:example.com ext:env OR ext:cfg OR ext:conf OR ext:ini 575 ``` 576 577 <https://www.google.com/search?q=site%3Aexample.com+ext%3Aenv...> 578 ~~~ 579 580 Paste it straight into the recon section of a report — see [Documentation and Reporting](/sheets/pentest-workflow/documentation-and-reporting). 581 582 > [!tip] The authorisation gate is once per machine 583 > The acknowledgement is stored in `$XDG_STATE_HOME/dorkforge/ack.json` (`~/.local/state/dorkforge/` by default) along with the scope reference you gave, which is then stamped into every session log. `--reset-ack` clears it.