daemon-sec-cheatsheet

The cheatsheet vault for operators: AD, enumeration, exploitation, priv-esc, web, DFIR
git clone https://git.daemon-sec.xyz/daemon-sec-cheatsheet.git
Log | Files | Refs | README | LICENSE

google-dorking.md (36739B)


      1 ---
      2 title: "Google Dorking"
      3 description: "Search-engine operators as a recon primitive: the full operator reference, the gotchas that silently break dorks, recipes by objective, cross-engine translation, OPSEC, and defending your own estate against them."
      4 category: enumeration
      5 tags: [enumeration, osint, recon, google-dorking, dorks, passive, search-operators, ghdb]
      6 tools: [Google, Bing, DuckDuckGo, Yandex, GitHub, Shodan, dorkforge]
      7 difficulty: beginner
      8 updated: "2026-09-26"
      9 ---
     10 
     11 # Google Dorking
     12 
     13 Dorking is using a search engine's own query language to find things its index already holds but nobody meant to publish: config files, backups, admin panels, stack traces, cloud buckets, private keys. No packet reaches the target — you are querying Google's copy of the web, which is why it opens [Stage 00 — Passive External Recon](/sheets/pentest-workflow/passive-external-recon) and why it survives every firewall, WAF and rate limit the target owns.
     14 
     15 The technique is old and the operators are few. What separates a useful dork from noise is knowing which operators still work, how they combine, and where each one silently fails.
     16 
     17 > [!danger] Indexed is not authorised
     18 > Every dork on this page returns *links*. Following one to retrieve a config file, enumerate a bucket, or log into a panel is **access**, and access needs written permission — a public URL is not consent, and "it was in Google" has never been a defence. Personal data you collect is regulated (GDPR and equivalents) whether or not the engagement covers it. Use this against your own estate, a scope you hold in writing, or a lab you are entitled to.
     19 
     20 ---
     21 
     22 ## 1. Why it works — you are searching the index, not the site
     23 
     24 A crawler finds a URL, fetches it, and the indexer stores the text, title, URL and file type. Everything dorking does is filter that stored record. Three consequences matter:
     25 
     26 | Fact | Consequence for recon |
     27 |---|---|
     28 | `robots.txt` is a **crawl** directive, not an access control | A `Disallow:` line stops polite crawlers fetching a path — it does not stop the path being served, and it publishes a list of the paths the owner considers sensitive. Fetch `$DOMAIN/robots.txt` first; it is a free sitemap of the interesting bits. |
     29 | Only `noindex` removes a page from results | `X-Robots-Tag: noindex` or a `<meta name="robots" content="noindex">` tag is the actual removal mechanism, and it only works if the crawler is *allowed in* to see it. A `Disallow`ed page that is linked from elsewhere can still appear. |
     30 | The index outlives the file | Deleting a file does not delete the record. Titles and snippets persist for weeks; archives persist indefinitely. A dork that returns a 404 still tells you the file existed and what it was called. |
     31 | Anything linked once is crawlable | A "secret" URL pasted into a public ticket, a README, a Slack export or a CDN log gets crawled. This is how private buckets and staging hosts end up indexed. |
     32 
     33 The practical reading: **the index is a historical record of what the target served, not a snapshot of what it serves now.** Treat every hit as a lead with a timestamp, not a live finding.
     34 
     35 ---
     36 
     37 ## 2. Anatomy of a dork
     38 
     39 ```text
     40 site:example.com ext:sql intext:"CREATE TABLE" -site:docs.example.com
     41 └──── scope ───┘ └─ type ┘ └───── content ────┘ └────── exclusion ─────┘
     42 ```
     43 
     44 Four moving parts, and only the first is mandatory in practice:
     45 
     46 - **Scope** — which hosts count (`site:`).
     47 - **Type** — what kind of document (`ext:` / `filetype:`).
     48 - **Content** — a literal that only the interesting version of the file contains (`intext:`, `intitle:`, `inurl:`).
     49 - **Exclusion** — what to subtract once the first pass is noisy (`-site:`, `-inurl:`).
     50 
     51 Terms and operators are **AND**-ed by default. There is no `AND` keyword — writing one searches for the word "and".
     52 
     53 > [!tip] Build them in that order
     54 > Scope first, then type, then narrow with content, then subtract noise. Starting from a long stacked dork you copied off a list gives you zero results and no idea which clause killed it. Start broad, add one operator at a time, and watch the result count move.
     55 
     56 ---
     57 
     58 ## 3. Operator reference
     59 
     60 ### Scope
     61 
     62 | Operator | Does | Example |
     63 |---|---|---|
     64 | `site:` | Restrict to a host. Matches the host as a **suffix**, so it covers subdomains. | `site:example.com` |
     65 | `site:*.` | Force the subdomain tree. Pair with a negative to drop the ones you know. | `site:*.example.com -site:www.example.com` |
     66 | `-site:` | Subtract a host. The main tool for forcing new subdomains to the surface. | `-site:blog.example.com` |
     67 | `related:` | Sites Google considers similar. **Officially unsupported** — usually returns nothing. | `related:example.com` |
     68 
     69 ### Content position
     70 
     71 | Operator | Matches in | Notes |
     72 |---|---|---|
     73 | `intitle:` | The `<title>` element | Highest-signal field for fingerprinting an app by its shipped default title. |
     74 | `inurl:` | The whole URL, **including the query string** | This is why it finds parameter names: `inurl:redirect=`. Substring match. |
     75 | `intext:` | The page body | Forces a term into the body instead of letting it match the title or an inbound link. |
     76 | `inanchor:` | Inbound link text | Thin and unreliable now that link signals are de-emphasised. |
     77 | `allintitle:` `allinurl:` `allintext:` | Same fields, all following terms | **Do not use.** See the gotchas below. |
     78 
     79 ### File type
     80 
     81 | Operator | Does | Notes |
     82 |---|---|---|
     83 | `filetype:` | One file extension | Matches the extension Google recorded, not the `Content-Type` header. |
     84 | `ext:` | Identical to `filetype:` | Shorter, so stacked dorks stay readable. Both work on Google, Bing and Brave. |
     85 
     86 One extension per operator — stack them with `OR` to cover a family:
     87 
     88 ```text
     89 site:$DOMAIN ext:sql OR ext:bak OR ext:old OR ext:backup
     90 ```
     91 
     92 Google indexes a fixed set of document types (`pdf`, `doc`/`docx`, `xls`/`xlsx`, `ppt`/`pptx`, `rtf`, `ps`, `kml`, `txt` and a handful more). A `ext:env` dork works **only because such files were served as plain text and indexed as text** — which is precisely what makes a hit worth reporting.
     93 
     94 ### Logic and grouping
     95 
     96 | Syntax | Does | Trap |
     97 |---|---|---|
     98 | `"phrase"` | Exact match, in order, no stemming or synonyms | Any literal artefact — a banner, an error string, a key header — belongs in quotes. |
     99 | `OR` or `\|` | Disjunction | **Must be uppercase.** Lowercase `or` is a stop word and is silently ignored. |
    100 | `-term` | Exclude | No space after the hyphen. Works on operators too: `-site:`, `-inurl:`. |
    101 | `( )` | Grouping | Required whenever you mix `OR` with AND-ed terms, or precedence bites you. |
    102 | `*` | Whole-token wildcard | Only meaningful **inside quotes**. It is not a substring wildcard — `adm*n` does not work. |
    103 | `AROUND(n)` | The two terms within *n* words | Uppercase, no space before the bracket. Tighter than AND, looser than a phrase. |
    104 
    105 ```text
    106 site:$DOMAIN (ext:sql OR ext:bak) intext:password        # grouped correctly
    107 site:$DOMAIN ext:sql OR ext:bak intext:password          # ambiguous — do not
    108 password AROUND(3) database site:$DOMAIN                 # proximity
    109 "confidential * report" site:$DOMAIN                     # token wildcard
    110 ```
    111 
    112 ### Time and number
    113 
    114 | Syntax | Does | Caveat |
    115 |---|---|---|
    116 | `before:YYYY-MM-DD` | Documents Google dates before that day | Filters on Google's *estimate* of the document date, which is frequently wrong for pages without date markup. |
    117 | `after:YYYY-MM-DD` | Documents Google dates after that day | Same caveat. Good for narrowing; never good enough to prove publication date. |
    118 | `1000..2000` | Any number in the range | Two dots. Legal inside other operators: `inurl:id=1..100`. |
    119 
    120 ### Retired — and what replaced them
    121 
    122 > [!warning] These appear in every dork list on the internet and none of them work
    123 > | Operator | Status | Use instead |
    124 > |---|---|---|
    125 > | `cache:` | Removed by Google in **September 2024**, along with the cached-page links | [web.archive.org](https://web.archive.org), [archive.today](https://archive.today) |
    126 > | `link:` | Deprecated **2017**; now a silent no-op that degrades to a keyword search | Search Console (for sites you own) or a commercial backlink tool |
    127 > | `info:` | Retired; folded into ordinary search | Just search the bare URL |
    128 > | `+term` | Removed **2011** | Quote the term instead: `"term"` |
    129 > | `daterange:` | Julian-date legacy syntax, unreliable | `before:` / `after:` |
    130 >
    131 > A retired operator does not error. Google drops it and searches the remaining words, so your dork returns plausible-looking garbage. This is the single most common reason a copied dork "stops working".
    132 
    133 ---
    134 
    135 ## 4. The gotchas that actually break dorks
    136 
    137 > [!example] Read this section twice — it is most of the skill
    138 > | Trap | What happens | Fix |
    139 > |---|---|---|
    140 > | Lowercase `or` | Treated as a stop word, dropped. Your disjunction became an AND. | Always uppercase `OR`. |
    141 > | `allintitle:` / `allinurl:` / `allintext:` | They consume **the rest of the query** — every operator after them is ignored, including `site:`. Your scoped dork quietly went global. | Stack the singular form instead: `intitle:a intitle:b`. |
    142 > | Space after the colon | `site: example.com` is parsed as a bare `site:` plus the keyword `example.com`. | Never put a space after an operator's colon. |
    143 > | Unquoted multi-word value | `intitle:index of` means `intitle:index` AND the word `of`. | Quote it: `intitle:"index of"`. |
    144 > | `*` outside quotes | Ignored entirely. | Only use it inside a quoted phrase. |
    145 > | `site:` is a suffix match | `site:example.com` also returns `notexample.com`-style matches on some engines and always includes every subdomain. | Add `-site:` exclusions, or use Yandex's stricter `host:`. |
    146 > | Very long stacked dorks | Google historically truncated long queries, and still degrades them — the tail of a 15-operator dork may never be evaluated. | Split into two or three narrower dorks. |
    147 > | Personalised results | Signed in, Google reorders and filters results for *you*. Two operators can look broken because your profile suppressed the hits. | Search signed out, in a clean profile. |
    148 > | Result count is a lie | The "About N results" figure is an estimate and swings wildly. | Page to the end. The real count is where results stop. |
    149 > | Verbatim mode | Google rewrites and expands queries by default. | Append `&tbs=li:1` to the results URL, or use Tools → All results → Verbatim. |
    150 
    151 > [!tip] Two switches worth memorising
    152 > `&num=100` on the results URL returns 100 per page instead of 10. `&filter=0` disables the "omitted similar results" collapse, which routinely hides the interesting duplicates on a large site.
    153 
    154 ---
    155 
    156 ## 5. Recipes by objective
    157 
    158 Every dork below is written in the Google dialect with `$DOMAIN`, `$ORG` and `$KEYWORD` as placeholders. Section 6 covers translating them.
    159 
    160 ### Exposed files and configs
    161 
    162 The highest-yield category in the whole practice.
    163 
    164 | Dork | Finds |
    165 |---|---|
    166 | `site:$DOMAIN ext:env OR ext:cfg OR ext:conf OR ext:ini` | Application config. A served `.env` is plaintext credentials: database URI, mail password, cloud keys, app secret. |
    167 | `site:$DOMAIN ext:bak OR ext:old OR ext:backup OR ext:swp` | Editor and deploy leftovers. `index.php.bak` is served as text instead of executed — that is the source code. |
    168 | `site:$DOMAIN ext:log OR ext:txt inurl:log` | Logs: internal paths, usernames, session tokens, stack traces. |
    169 | `site:$DOMAIN ext:yml OR ext:yaml OR ext:toml inurl:config` | Compose files, CI definitions, `appsettings`, secrets baked into a values file. |
    170 | `site:$DOMAIN ext:pem OR ext:key OR ext:ppk OR ext:p12 OR ext:pfx` | Private keys and cert bundles. Any hit is a critical finding on its own. |
    171 | `site:$DOMAIN inurl:.git OR inurl:.svn OR inurl:.hg` | Exposed VCS metadata — a readable `.git/` means the whole repo and its history. |
    172 | `site:$DOMAIN inurl:wp-config OR inurl:configuration.php OR inurl:settings.py` | Framework config by canonical filename. |
    173 
    174 ### Login and admin panels
    175 
    176 | Dork | Finds |
    177 |---|---|
    178 | `site:$DOMAIN inurl:admin OR inurl:administrator OR inurl:adminpanel` | The obvious paths. Still hits constantly. |
    179 | `site:$DOMAIN intitle:"login" OR intitle:"sign in"` | Title-based, so it catches panels on paths you would never guess. |
    180 | `site:$DOMAIN inurl:phpmyadmin OR inurl:adminer OR inurl:pma` | Exposed database consoles, frequently on vendor defaults. |
    181 | `site:$DOMAIN intitle:"Dashboard" (inurl:jenkins OR inurl:grafana OR inurl:kibana)` | CI and observability — build secrets, internal hostnames, production telemetry. |
    182 | `site:$DOMAIN inurl:/manager/html OR inurl:/jmx-console OR inurl:/axis2` | Java app-server managers. Tomcat manager with default creds is a WAR deploy from RCE. |
    183 | `site:$DOMAIN inurl:owa OR inurl:/rdweb OR inurl:citrix OR inurl:vpn` | Remote-access front doors — the surface spraying actually targets. |
    184 
    185 ### Open directory listings
    186 
    187 | Dork | Finds |
    188 |---|---|
    189 | `site:$DOMAIN intitle:"index of"` | The canonical dork. Apache, nginx and IIS all emit a title starting this way. |
    190 | `site:$DOMAIN intitle:"index of" (backup OR bak OR old OR archive)` | Listings that name themselves as a dumping ground. |
    191 | `site:$DOMAIN intitle:"index of" "parent directory" (.sql OR .zip OR .tar.gz)` | Listings with a dump or archive sitting in them. |
    192 | `site:$DOMAIN "Directory Listing For" -inurl:html` | Tomcat and others whose listing wording differs from Apache's. |
    193 
    194 ### Cloud storage
    195 
    196 | Dork | Finds |
    197 |---|---|
    198 | `site:s3.amazonaws.com "$ORG"` | S3 buckets whose listing mentions the organisation. |
    199 | `site:$DOMAIN inurl:s3.amazonaws.com` | Pages on the target that link into S3 — the fastest way to learn the bucket naming convention. |
    200 | `inurl:blob.core.windows.net "$ORG"` | Azure Blob containers. |
    201 | `inurl:storage.googleapis.com "$ORG"` | Google Cloud Storage objects. |
    202 | `site:$DOMAIN inurl:r2.dev OR inurl:digitaloceanspaces.com` | Smaller providers — less scrutiny, misconfigured more often. |
    203 | `"$ORG" (site:trello.com OR site:notion.site)` | Public boards and docs. Trello leaks credentials and architecture diagrams at a remarkable rate. |
    204 
    205 ### Credentials, keys and tokens
    206 
    207 | Dork | Finds |
    208 |---|---|
    209 | `site:$DOMAIN intext:"BEGIN RSA PRIVATE KEY"` | Key headers pasted into a wiki, a ticket, a README. |
    210 | `site:$DOMAIN intext:"AKIA"` | AWS access key IDs all begin `AKIA` — a distinctive literal that rarely false-positives. |
    211 | `"$ORG" ("xoxb-" OR "ghp_" OR "sk_live_")` | Slack, GitHub and Stripe token prefixes. Structurally unique, so a hit is almost never noise. |
    212 | `site:$DOMAIN (ext:env OR ext:yml) intext:PASSWORD` | Config files with a password assignment. |
    213 | `site:pastebin.com OR site:justpaste.it "$DOMAIN"` | Paste sites — where dumps, configs and credential lists get parked. |
    214 
    215 > [!danger] A live key is a finding, not a login
    216 > Do not authenticate with anything you find. Record the URL, report it immediately so it can be rotated, and treat your own notes as sensitive material for the rest of the engagement. See [Credential Hunting](/sheets/enumeration/credential-hunting), [Gitleaks](/sheets/enumeration/2-4-cheatsheet-gitleaks) and [TruffleHog](/sheets/enumeration/2-5-cheatsheet-trufflehog) for the systematic version of this.
    217 
    218 ### Documents and metadata
    219 
    220 | Dork | Finds |
    221 |---|---|
    222 | `site:$DOMAIN ext:pdf OR ext:docx OR ext:xlsx OR ext:pptx` | The corpus. Pull it, then run `exiftool` over the lot. |
    223 | `site:$DOMAIN ext:pdf (confidential OR internal OR "not for distribution")` | Documents labelled restricted that are nonetheless public. |
    224 | `site:$DOMAIN ext:xlsx (employee OR salary OR roster)` | Spreadsheets — the format people paste staff lists into. |
    225 | `site:$DOMAIN ext:vsd OR ext:vsdx OR ext:drawio` | Network and architecture diagrams. |
    226 
    227 ```bash
    228 # the half of this that isn't the dork: metadata gives you usernames and internal paths
    229 exiftool -Author -Creator -Producer -Company *.pdf | sort -u
    230 exiftool -a -G1 -s report.docx | grep -Ei 'author|company|template|lastmodifiedby'
    231 ```
    232 
    233 ### Error messages and stack traces
    234 
    235 | Dork | Finds |
    236 |---|---|
    237 | `site:$DOMAIN intext:"You have an error in your SQL syntax"` | The literal MySQL parser message — confirms injectable input. |
    238 | `site:$DOMAIN intext:"Traceback (most recent call last)"` | Python tracebacks. Django debug pages additionally dump settings and environment. |
    239 | `site:$DOMAIN intitle:"Whoops, looks like something went wrong"` | Laravel's debug page, which shows environment variables verbatim. |
    240 | `site:$DOMAIN intext:"Fatal error" OR intext:"Uncaught exception"` | PHP and Java fatals — usually carry the absolute filesystem path. |
    241 | `site:$DOMAIN intext:"Server Error in" intext:"Stack Trace"` | ASP.NET yellow-screen-of-death with tracing left on. |
    242 | `site:$DOMAIN intitle:"phpinfo()"` | Full build config, loaded modules, absolute paths, sometimes environment secrets. |
    243 
    244 ### Databases, backups and dumps
    245 
    246 | Dork | Finds |
    247 |---|---|
    248 | `site:$DOMAIN ext:sql intext:"INSERT INTO"` | SQL dumps confirmed by their own statement syntax, not just an extension. |
    249 | `site:$DOMAIN ext:sql intext:"CREATE TABLE" intext:password` | Dumps that include a credentials table. |
    250 | `site:$DOMAIN (ext:zip OR ext:tar OR ext:gz) (backup OR dump OR export)` | Archived backups in the webroot. |
    251 | `site:$DOMAIN inurl:backup OR inurl:dump OR inurl:export` | Path-based — catches the directory even when the file was not indexed. |
    252 
    253 ### Injectable parameters
    254 
    255 A target list for later, authorised testing. Nothing here is a finding by itself.
    256 
    257 | Dork | Vulnerability class |
    258 |---|---|
    259 | `site:$DOMAIN inurl:id= OR inurl:pid= OR inurl:cat=` | SQL injection, IDOR |
    260 | `site:$DOMAIN inurl:file= OR inurl:page= OR inurl:path=` | LFI, path traversal — see [LFI](/sheets/enumeration/lfi) |
    261 | `site:$DOMAIN inurl:url= OR inurl:redirect= OR inurl:next=` | Open redirect, the front half of many SSRF chains |
    262 | `site:$DOMAIN inurl:cmd= OR inurl:exec= OR inurl:query=` | Command injection, reflected XSS |
    263 | `site:$DOMAIN inurl:debug=true OR inurl:admin=1` | Flags that flip the app into a verbose or privileged mode |
    264 
    265 ### People, emails and usernames
    266 
    267 | Dork | Finds |
    268 |---|---|
    269 | `site:linkedin.com/in "$ORG"` | Staff profiles — names plus role, which is what derives the username scheme. |
    270 | `site:$DOMAIN intext:"@$DOMAIN"` | The org's own pages leaking its address format (`first.last@`, `flast@`). |
    271 | `"$ORG" (site:github.com OR site:gitlab.com) intext:"@$DOMAIN"` | Developers using a work address in commits or profiles. |
    272 | `"$ORG" (site:stackoverflow.com OR site:serverfault.com)` | Engineers asking about their own stack, pasting real config with real hostnames. |
    273 | `"$ORG" ("we use" OR "powered by") -site:$DOMAIN` | Tech stack from job ads, case studies, conference talks. |
    274 
    275 > [!warning] Personal data is regulated independently of your scope
    276 > Collect the minimum that supports the finding, store it with the engagement material, and delete it when the engagement closes.
    277 
    278 ### Subdomains and forgotten hosts
    279 
    280 | Dork | Finds |
    281 |---|---|
    282 | `site:*.$DOMAIN -site:www.$DOMAIN` | The core subdomain dork. |
    283 | `site:*.$DOMAIN (dev OR staging OR test OR uat OR qa)` | Non-production — weaker credentials, debug on, real data. |
    284 | `site:*.$DOMAIN (vpn OR remote OR portal OR intranet)` | Access infrastructure and internal-facing hosts. |
    285 | `"$DOMAIN" -site:$DOMAIN` | Third-party references: partners, status pages, monitoring, archived copies. |
    286 
    287 Dorking is the *weakest* of the subdomain sources — use it to catch what certificate transparency and DNS missed, not as the primary pass. See [Amass](/sheets/enumeration/amass) and [Stage 00](/sheets/pentest-workflow/passive-external-recon) for the real enumeration.
    288 
    289 ### Archived and removed pages
    290 
    291 | Dork | Finds |
    292 |---|---|
    293 | `site:web.archive.org "$DOMAIN"` | Wayback captures Google indexed. |
    294 | `site:archive.today "$DOMAIN"` | The other archive, which captures pages Wayback refuses. |
    295 | `site:$DOMAIN before:2022-01-01 "$KEYWORD"` | Older material still on the live site. |
    296 
    297 ```bash
    298 # the archive's own API beats dorking it — every captured URL, deduplicated
    299 curl -s "http://web.archive.org/cdx/search/cdx?url=$DOMAIN*&output=text&fl=original&collapse=urlkey" \
    300   | sort -u > wayback-urls.txt
    301 # pull the ones that look like config or secrets out of the pile
    302 grep -Ei '\.(env|sql|bak|old|log|yml|ini|conf)($|\?)' wayback-urls.txt
    303 ```
    304 
    305 ---
    306 
    307 ## 6. Cross-engine translation
    308 
    309 Google is the reference dialect, not the only one. Yandex and Bing routinely index hosts Google never crawled, so a subdomain dork is worth running through at least two engines.
    310 
    311 | Intent | Google | Bing | DuckDuckGo | Yandex | Brave |
    312 |---|---|---|---|---|---|
    313 | Scope to host | `site:` | `site:` | `site:` | `site:` / `host:` / `rhost:` | `site:` |
    314 | File type | `filetype:` `ext:` | `filetype:` `ext:` | `filetype:` | `mime:` | `filetype:` |
    315 | In title | `intitle:` | `intitle:` | `intitle:` | `title:` | `intitle:` |
    316 | In URL | `inurl:` | `inurl:` | `inurl:` | `url:` | `inurl:` |
    317 | In body | `intext:` | `inbody:` | — | *(default)* | — |
    318 | Exact phrase | `"..."` | `"..."` | `"..."` | `"..."` | `"..."` |
    319 | Exclude | `-term` | `-term` | `-term` | `-term` | `-term` |
    320 | Either/or | `OR` | `OR` | `OR` | `OR` | `OR` |
    321 | Proximity | `AROUND(n)` | — | — | `/+n` | — |
    322 | Date bound | `before:` `after:` | *(UI filter)* | *(UI filter)* | `date:` | *(UI filter)* |
    323 
    324 **Engine-native operators with no Google equivalent — worth the detour on their own:**
    325 
    326 | Engine | Operator | Why you would bother |
    327 |---|---|---|
    328 | Bing | `ip:203.0.113.10` | Every site Bing indexed on that address — a free reverse-IP lookup, and often the rest of a shared estate behind one origin. |
    329 | Bing | `contains:sql` | Pages that **link to** a `.sql` file rather than the file itself — catches the index page of a backup directory that was never indexed. |
    330 | Bing | `url:example.com/path` | Tests whether one exact URL is in the index. |
    331 | Yandex | `rhost:com.example.*` | Reversed-domain match across the whole tree — stricter and more predictable than Google's suffix matching. |
    332 | Yandex | `mime:pdf` | Yandex's `filetype:`, keyed on recorded MIME type. |
    333 | DuckDuckGo | `!bangs` (`!gh`, `!so`) | Not operators — redirects to another site's own search. Handy, but the query leaves DDG. |
    334 
    335 > [!note] Brave and DuckDuckGo are the low-friction options
    336 > Neither requires an account and neither personalises results, which makes them the right place to start from a clean browser. Their operator support is genuinely narrower than Google's and less documented — when a dork returns nothing there, re-run it on Google before concluding the target is clean.
    337 
    338 ---
    339 
    340 ## 7. Other dialects worth knowing
    341 
    342 ### GitHub code search
    343 
    344 A different language entirely, and the single richest source of leaked secrets. Requires being signed in — use a throwaway account.
    345 
    346 | Operator | Example | Notes |
    347 |---|---|---|
    348 | `org:` | `org:$ORG path:.env` | Scope to an organisation's public repos. |
    349 | `repo:` | `repo:owner/name "password"` | One repository. |
    350 | `path:` | `path:.github/workflows` | Replaced the old `filename:`. Globs: `path:**/config/*.yml`. |
    351 | `language:` | `"$DOMAIN" language:yaml` | Narrows a noisy content search. |
    352 | `symbol:` | `symbol:connectDatabase` | Definitions rather than mentions. |
    353 
    354 ```text
    355 org:$ORG path:.env
    356 org:$ORG "BEGIN RSA PRIVATE KEY"
    357 org:$ORG path:.github/workflows
    358 "$DOMAIN" ("192.168." OR "10.0." OR ".local")
    359 "$DOMAIN" "password" language:yaml
    360 ```
    361 
    362 The last two are the interesting ones: they find *other people's* repos that leak the target's internal addressing and credentials — contractors, ex-employees, integration partners.
    363 
    364 ### Shodan and Censys
    365 
    366 These index service banners, not web pages, so the web operators do not apply at all. Covered properly in [Shodan](/sheets/enumeration/shodan); the translation for a dorking reflex:
    367 
    368 | You want | Shodan |
    369 |---|---|
    370 | Hosts presenting the org's certificate | `ssl.cert.subject.cn:$DOMAIN` |
    371 | Pages referencing the domain | `http.html:"$DOMAIN"` |
    372 | A specific page title | `http.title:"index of"` |
    373 | The org's netblocks | `org:"$ORG"` or `net:203.0.113.0/24` |
    374 
    375 The origin-IP hunt behind a CDN is `ssl.cert.subject.cn:` — the origin still presents the real certificate.
    376 
    377 ---
    378 
    379 ## 8. The Google Hacking Database
    380 
    381 The [GHDB](https://www.exploit-db.com/google-hacking-database) is Exploit-DB's curated archive of dorks, started by Johnny Long, now thousands of entries across categories: files containing passwords, sensitive directories, vulnerable servers, error messages, footholds.
    382 
    383 How to use it without wasting an afternoon:
    384 
    385 - **Filter by date.** Most of the archive is a decade old and targets software nobody runs. Sort newest first.
    386 - **Read it as a pattern library, not a copy-paste list.** The value of `intitle:"index of" "service.pwd"` is not that exact string — it is that vendor-default filenames are the way in. Adapt the shape to the stack you actually found.
    387 - **Scope everything.** Almost every GHDB entry is unscoped and returns strangers' infrastructure. Add `site:$DOMAIN` before you run it, every time.
    388 
    389 ---
    390 
    391 ## 9. OPSEC
    392 
    393 > [!danger] You are the one being logged
    394 > Dorking sends nothing to the target — but it sends everything to the search engine, tied to your account, your IP and your browser fingerprint. The queries themselves describe your intent with unusual clarity.
    395 
    396 | Rule | Why |
    397 |---|---|
    398 | Never dork signed in to a real account | The query history is retained, attributable, and personalises your results into uselessness. Use a clean profile or a private window. |
    399 | Use a throwaway for GitHub code search | It is the one dialect that mandates an account. Do not attach your identity to a secrets hunt. |
    400 | Expect CAPTCHAs, and slow down | A burst of operator-heavy queries from one IP trips Google's automation detection. Once you are CAPTCHA'd, that IP is degraded for hours. |
    401 | Do not click through to live secrets | Fetching the `.env` is a request to the target — logged, attributable, and outside "passive" recon. The snippet in the result usually contains enough to report. |
    402 | Treat your notes as sensitive | A session log full of working dorks and hit URLs is a map of the target's exposure. Store it with the engagement, encrypt it, delete it on close. |
    403 | Report live keys immediately | Rotation is time-sensitive. Do not sit on a working credential until the report. |
    404 
    405 ---
    406 
    407 ## 10. Automation, rate limits and the ToS question
    408 
    409 Scripted scraping of Google results violates its Terms of Service and fails in practice: the automation detection is good, the HTML changes, and the CAPTCHA is unsolvable at scale without paying someone to solve it. Every "google dork scanner" on GitHub has the same lifecycle — works for a week, then returns empty pages forever.
    410 
    411 The options that actually work:
    412 
    413 | Approach | Reality |
    414 |---|---|
    415 | **Do it by hand** | Slow, unblockable, and you read each result properly. For a single engagement this is genuinely the right answer. |
    416 | **Programmable Search Engine JSON API** | Google's sanctioned path. Free tier is ~100 queries/day, then paid. Supports most operators. Results are scoped to a custom engine you configure, which can be set to search the whole web. |
    417 | **A commercial SERP API** | Pays someone else to absorb the blocking. Fine for volume; check that sending client data to a third party is acceptable under your engagement's confidentiality terms. |
    418 | **Shodan / Censys APIs** | Purpose-built, properly documented, generous limits. If the thing you want is a service rather than a document, stop dorking and query these. |
    419 | **Bing Search APIs** | Microsoft has been winding these down — verify current availability before building anything on them. |
    420 
    421 A tool that *builds* dorks and hands them to your browser has none of these problems, because you are still the one searching. That is what [dorkforge](#13-dorkforge--the-companion-script) does.
    422 
    423 ---
    424 
    425 ## 11. Blue team — defending your own estate
    426 
    427 Dorking is trivially turned around: run it against yourself, on a schedule.
    428 
    429 | Control | What it actually does |
    430 |---|---|
    431 | `X-Robots-Tag: noindex` header, or `<meta name="robots" content="noindex">` | **The** removal mechanism. Note the trap: a path blocked in `robots.txt` cannot be crawled, so the crawler never sees the `noindex` — the page can stay indexed forever. Allow the crawl, serve the `noindex`, then block it once it has dropped out. |
    432 | Search Console removal tool | Fast takedown for a site you own (~6 months), which buys time while you fix the underlying exposure. It is a suppression, not a fix. |
    433 | `robots.txt` | A crawl hint, and a published list of your sensitive paths. Never treat it as access control; assume attackers read it first. |
    434 | Disable autoindex | `Options -Indexes` (Apache), `autoindex off` (nginx). Kills the entire `intitle:"index of"` class. |
    435 | Block dotfiles and backup extensions at the edge | Return 404 for `\.(env\|git\|bak\|old\|sql\|swp)$` before the request reaches the app. |
    436 | Turn off debug in production | `APP_DEBUG=false`, `customErrors="On"`, `DEBUG = False`. Removes the entire error-message class in one change. |
    437 | Bucket ACLs and public-access blocks | Cloud storage exposure is a configuration problem; no search-engine control fixes it. |
    438 | Continuous secret scanning | Gitleaks/TruffleHog in CI plus GitHub push protection catches the commit before it is ever indexed. |
    439 | Dork yourself on a schedule | The blue-team use of this whole page. Automate the handful of dorks that matter for your estate and alert on new hits. |
    440 
    441 ```bash
    442 # a minimal self-audit, run monthly against your own domain
    443 for d in 'ext:env' 'ext:sql' 'ext:bak' 'intitle:"index of"' 'intext:"BEGIN RSA PRIVATE KEY"'; do
    444   printf '%s\n' "site:$DOMAIN $d"
    445 done
    446 # feed them to the Programmable Search JSON API and diff the hit list against last month's
    447 ```
    448 
    449 ---
    450 
    451 ## 12. Practising safely
    452 
    453 - **Your own domains.** The only fully unambiguous target, and the one where findings are actionable.
    454 - **`google-gruyere.appspot.com`**, **`testphp.vulnweb.com`**, **`demo.testfire.net`** — deliberately vulnerable hosts published for training. Verify the terms on each before touching it.
    455 - **The GHDB, read but not run.** Study the shapes; run them scoped to yourself.
    456 - **CTF and lab platforms** with an OSINT category, where the scope is explicit.
    457 
    458 > [!warning] Unscoped dorks return strangers' infrastructure
    459 > A dork like `intitle:"webcamXP"` returns hardware belonging to people who never consented to anything. Looking at the result list is one thing; connecting to a device is unauthorised access in essentially every jurisdiction. Scope with `site:` or do not run it.
    460 
    461 ---
    462 
    463 ## 13. dorkforge — the companion script
    464 
    465 **[dorkforge.py](/downloads/enumeration/dorkforge.py)** ([SHA-256](/downloads/enumeration/dorkforge.py.sha256)) is a single-file interactive workbench for everything on this page: it asks what you are hunting for, builds the dork, **explains every operator it used**, translates it for the engine you picked, and hands it to your clipboard, your browser, or an engagement log.
    466 
    467 It ships the full library from section 5 — **14 objectives, 81 recipes, 32 operators** — and it sends no traffic to any target. The searching still happens in your browser, under your account, from your IP, which is why it opens with the authorisation gate and records a scope reference in every session log.
    468 
    469 ### Run it
    470 
    471 It uses [PEP 723](https://peps.python.org/pep-0723/) inline metadata, so `uv` resolves `rich` and `questionary` into a throwaway environment on first run. Nothing to install, nothing to clean up.
    472 
    473 ```bash
    474 # download + verify
    475 curl -sO https://cheatsheet.daemon-sec.xyz/downloads/enumeration/dorkforge.py
    476 curl -sO https://cheatsheet.daemon-sec.xyz/downloads/enumeration/dorkforge.py.sha256
    477 shasum -a 256 -c dorkforge.py.sha256     # expect: dorkforge.py: OK
    478 
    479 # run it
    480 uv run dorkforge.py
    481 ```
    482 
    483 Expected SHA-256:
    484 
    485 ```text
    486 3173df88ab3102ebc4999985fcb881298f3f51374170592e5ac95d0674210193
    487 ```
    488 
    489 Or make it executable and let the shebang do the work — `#!/usr/bin/env -S uv run --script`:
    490 
    491 ```bash
    492 chmod +x dorkforge.py
    493 ./dorkforge.py
    494 # keep it on PATH like the rest of your toolkit
    495 install -m 755 dorkforge.py ~/.local/bin/dorkforge
    496 ```
    497 
    498 > [!note] Don't have uv?
    499 > `curl -LsSf https://astral.sh/uv/install.sh | sh` (or `brew install uv`). Failing that, `pip install rich questionary` and run it with plain `python3` — the script degrades to that path and tells you so.
    500 
    501 ### The flow
    502 
    503 ```text
    504   ▪ DÆMON//SEC
    505   dorkforge v1.0.0    search-operator workbench
    506 ────────────────────────────────────────────────────────────
    507 
    508 ? What are you hunting for   Exposed files & configs
    509 ? Target domain: example.com
    510 ? Which engine? Google
    511 ? Which dork?  site:example.com ext:env OR ext:cfg OR ext:conf OR ext:ini
    512 
    513 ┌─   Google ─────────────────────────────────────────────────┐
    514 │                                                            │
    515 │  site:example.com ext:env OR ext:cfg OR ext:conf OR ext:ini│
    516 │                                                            │
    517 │  Application config. A served .env is credentials in       │
    518 │  plaintext - database URI, mail password, cloud keys.      │
    519 │                                                            │
    520 └────────────────────────────────────────────────────────────┘
    521 
    522  OPERATOR   WHAT IT DOES                               GOOGLE
    523  ──────────────────────────────────────────────────────────────
    524  site:      Only return pages whose host matches.        ✓
    525  ext:       Synonym for filetype: on Google and Bing.    ✓
    526  OR         Match either side. Must be uppercase.        ✓
    527 
    528 ? Now what?  Copy to clipboard / Open in browser / Save to log
    529 ```
    530 
    531 Pick a non-Google engine and it rewrites the dork and tells you what changed:
    532 
    533 ```text
    534 $ uv run dorkforge.py --objective listing --domain example.com --engine yandex
    535 
    536   site:example.com title:"index of"
    537     dialect notes
    538      ▪ intitle:  rewritten to title: for Yandex
    539 ```
    540 
    541 ### Usage
    542 
    543 | Command | Does |
    544 |---|---|
    545 | `uv run dorkforge.py` | Interactive. The main path. |
    546 | `dorkforge --session engagement.md` | Append every dork you build to a markdown log, with the URL, the reasoning and the translation notes. |
    547 | `dorkforge --objective files --domain example.com` | Non-interactive: print every recipe for one objective. |
    548 | `dorkforge --objective code --org acme --engine github` | Same, in GitHub's dialect. |
    549 | `dorkforge --list` | The whole library — 14 objectives, 81 recipes — to a pager or a file. |
    550 | `dorkforge --operators` | The operator reference table, with live / unreliable / retired status. |
    551 | `dorkforge --explain site:` | The long form on any one operator, plus which engines support it. |
    552 | `dorkforge --doctor` | Check clipboard helper, browser, colour support and state directory. |
    553 | `dorkforge --ascii` | ASCII markers instead of Nerd Font glyphs. |
    554 | `dorkforge --reset-ack` | Forget the authorisation acknowledgement and ask again. |
    555 
    556 **Flags:** `--domain` `--org` `--keyword` `--ext` pre-fill the scope prompts. `--engine` picks the dialect (`google` `bing` `ddg` `yandex` `brave` `github` `shodan`). `--no-color` for piping.
    557 
    558 ### The session log
    559 
    560 `--session` writes engagement-ready markdown as you work — the dork, why it was built, the live URL, and anything that did not survive translation:
    561 
    562 ~~~markdown
    563 # dorkforge session
    564 
    565 - **Started:** 2026-09-26 14:21 BST
    566 - **Authorised scope:** ACME-2026-114
    567 - **Tool:** dorkforge 1.0.0
    568 
    569 ## 1. Google
    570 
    571 Application config. A served .env is credentials in plaintext.
    572 
    573 ```text
    574 site:example.com ext:env OR ext:cfg OR ext:conf OR ext:ini
    575 ```
    576 
    577 <https://www.google.com/search?q=site%3Aexample.com+ext%3Aenv...>
    578 ~~~
    579 
    580 Paste it straight into the recon section of a report — see [Documentation and Reporting](/sheets/pentest-workflow/documentation-and-reporting).
    581 
    582 > [!tip] The authorisation gate is once per machine
    583 > The acknowledgement is stored in `$XDG_STATE_HOME/dorkforge/ack.json` (`~/.local/state/dorkforge/` by default) along with the scope reference you gave, which is then stamped into every session log. `--reset-ack` clears it.