daemon-sec-cheatsheet

The cheatsheet vault for operators: AD, enumeration, exploitation, priv-esc, web, DFIR
git clone https://git.daemon-sec.xyz/daemon-sec-cheatsheet.git
Log | Files | Refs | README | LICENSE

commit 8ce68e70d359451b37def46c69061e23238de382
parent 678e012e70b52aba81298ac65dd92b1ba6556ff2
Author: $: DAΞMON <zer0sec.xp@icloud.com>
Date:   Sat, 26 Sep 2026 14:46:38 +0100

Add Google dorking cheat sheet and the dorkforge companion script

Expands the single dork table in passive-external-recon.md into a full
enumeration sheet: operator reference, the retired operators that silently
degrade a query (cache: pulled Sept 2024, link:, info:, +), the traps that
break dorks in practice, per-objective recipe tables, cross-engine
translation, OPSEC, and a blue-team section on defending an estate against
the same searches.

Ships dorkforge.py under public/downloads/enumeration/ - a single-file
interactive workbench for the same library. PEP 723 inline metadata so
`uv run dorkforge.py` resolves rich + questionary with nothing to install.
It builds a dork, explains every operator it used, translates it for the
target engine, and hands it to the clipboard, the browser or a markdown
session log. 14 objectives, 81 recipes, 32 operators, Rose Pine night
palette, and a one-time authorisation gate that records a scope reference
into every log.

No .sha256.asc link here: there is not one signature file in the repo, so
the pattern used elsewhere in the sheets resolves to a 404. Ships the file,
its .sha256 sidecar and a per-category SHA256SUMS instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Diffstat:
Mcontent-manifest.json | 28++++++++++++++++++++++++++++
Apublic/downloads/enumeration/SHA256SUMS | 1+
Apublic/downloads/enumeration/dorkforge.py | 1945+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Apublic/downloads/enumeration/dorkforge.py.sha256 | 1+
Asrc/content/sheets/enumeration/google-dorking.md | 583+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Msrc/content/sheets/pentest-workflow/passive-external-recon.md | 2++
Msrc/content/sheets/pentest-workflow/tool-index.md | 2+-
7 files changed, 2561 insertions(+), 1 deletion(-)

diff --git a/content-manifest.json b/content-manifest.json @@ -1258,6 +1258,34 @@ "difficulty": "intermediate", "action": "add", "note": "New sheet: v5 rewrote the CLI (OAM graph DB, engine subcommand, intel folded into enum) vs the v3/v4-era commands still shown in passive-external-recon.md." + }, + { + "source": "local", + "rel": "dorkforge 1.0.0", + "category": "enumeration", + "slug": "google-dorking", + "title": "Google Dorking", + "description": "Search-engine operators as a recon primitive \u2014 full operator reference, the gotchas that silently break dorks, recipes by objective, cross-engine translation, OPSEC, and blue-team defence.", + "tools": [ + "Google", + "Bing", + "DuckDuckGo", + "Yandex", + "GitHub", + "Shodan", + "dorkforge" + ], + "tags": [ + "enumeration", + "osint", + "recon", + "google-dorking", + "dorks", + "passive" + ], + "difficulty": "beginner", + "action": "add", + "note": "New sheet: expands the single dork table in passive-external-recon.md into a full reference, and ships the dorkforge.py companion script under public/downloads/enumeration/." } ] } \ No newline at end of file diff --git a/public/downloads/enumeration/SHA256SUMS b/public/downloads/enumeration/SHA256SUMS @@ -0,0 +1 @@ +3173df88ab3102ebc4999985fcb881298f3f51374170592e5ac95d0674210193 dorkforge.py diff --git a/public/downloads/enumeration/dorkforge.py b/public/downloads/enumeration/dorkforge.py @@ -0,0 +1,1945 @@ +#!/usr/bin/env -S uv run --script +# /// script +# requires-python = ">=3.11" +# dependencies = [ +# "rich>=13.7", +# "questionary>=2.0", +# ] +# /// +"""dorkforge - DAEMON//SEC search-operator workbench. + +An interactive coach for search-engine dorking. Pick what you are hunting +for, answer a couple of scope questions, and dorkforge builds the query, +explains every operator it used, translates it for the engine you picked, +and hands it to your clipboard, your browser, or an engagement log. + + uv run dorkforge.py + +Everything it produces is a *search query*. It sends no traffic to any +target and needs no API key. The searching still happens in your browser, +under your account, from your IP - which is exactly why the OPSEC notes +matter. + +Full cheatsheet: https://cheatsheet.daemon-sec.xyz/sheets/enumeration/google-dorking +""" + +from __future__ import annotations + +import argparse +import json +import os +import platform +import re +import shutil +import subprocess +import sys +import textwrap +import urllib.parse +import webbrowser +from dataclasses import dataclass, field +from datetime import datetime, timezone +from pathlib import Path + +try: + import questionary + from questionary import Choice, Style as QStyle + from rich import box + from rich.console import Console, Group + from rich.panel import Panel + from rich.rule import Rule + from rich.table import Table + from rich.text import Text + from rich.theme import Theme +except ModuleNotFoundError: # pragma: no cover - only hit when run bare + sys.exit( + "dorkforge needs rich + questionary.\n" + "Run it through uv so they are resolved for you:\n" + " uv run dorkforge.py\n" + "or install them yourself: pip install rich questionary" + ) + +VERSION = "1.0.0" +SHEET_URL = "https://cheatsheet.daemon-sec.xyz/sheets/enumeration/google-dorking" + +# --------------------------------------------------------------------------- +# Rose Pine (night) - the contrast-tuned DAEMON//SEC values, not stock upstream. +# `pine` here is #3e8fb0, lifted from the published #31748f, which sits at +# 3.3:1 on this background and is a fill colour, never ink. +# --------------------------------------------------------------------------- +BASE = "#191724" +SURFACE = "#1f1d2e" +OVERLAY = "#26233a" +MUTED = "#8b87a3" +SUBTLE = "#908caa" +TEXT = "#e0def4" +LOVE = "#eb6f92" +GOLD = "#f6c177" +ROSE = "#ebbcba" +PINE = "#3e8fb0" +FOAM = "#9ccfd8" +IRIS = "#c4a7e7" +HL_MED = "#403d52" + +THEME = Theme( + { + "df.text": TEXT, + "df.dim": SUBTLE, + "df.faint": MUTED, + "df.rule": HL_MED, + "df.love": LOVE, + "df.gold": GOLD, + "df.rose": ROSE, + "df.pine": PINE, + "df.foam": FOAM, + "df.iris": IRIS, + "df.op": f"bold {FOAM}", + "df.val": GOLD, + "df.kw": IRIS, + "df.warn": f"bold {GOLD}", + "df.bad": f"bold {LOVE}", + "df.ok": f"bold {FOAM}", + "df.eyebrow": f"bold {MUTED}", + "df.head": f"bold {TEXT}", + } +) + +QUESTIONARY_STYLE = QStyle( + [ + ("qmark", f"fg:{FOAM} bold"), + ("question", f"fg:{TEXT} bold"), + ("answer", f"fg:{GOLD} bold"), + ("pointer", f"fg:{LOVE} bold"), + ("highlighted", f"fg:{IRIS} bold"), + ("selected", f"fg:{FOAM}"), + ("separator", f"fg:{HL_MED}"), + ("instruction", f"fg:{MUTED}"), + ("text", f"fg:{TEXT}"), + ("disabled", f"fg:{MUTED} italic"), + ] +) + +# --- Nerd Font glyphs, with an --ascii fallback ---------------------------- +GLYPHS = { + "search": "", + "file": "", + "lock": "", + "cloud": "", + "db": "", + "bug": "", + "camera": "", + "code": "", + "users": "", + "folder": "", + "globe": "", + "book": "", + "bolt": "", + "check": "", + "cross": "", + "info": "", + "warn": "", + "copy": "", + "save": "", + "key": "", + "shield": "", + "arrow": "", + "dot": "▪", +} +ASCII_GLYPHS = { + "search": ">", "file": "f", "lock": "#", "cloud": "^", "db": "=", + "bug": "!", "camera": "o", "code": "<", "users": "&", "folder": "/", + "globe": "@", "book": "b", "bolt": "*", "check": "+", "cross": "x", + "info": "i", "warn": "!", "copy": "c", "save": "s", "key": "k", + "shield": "%", "arrow": "->", "dot": "-", +} +_G = dict(GLYPHS) + + +def g(name: str) -> str: + return _G.get(name, "") + + +# --------------------------------------------------------------------------- +# Operator reference +# --------------------------------------------------------------------------- +@dataclass(frozen=True) +class Operator: + key: str # canonical token, e.g. "site:" + label: str # human name + summary: str # one line, shown in the explain table + detail: str # the paragraph shown by `--explain` + example: str + status: str = "live" # live | flaky | retired + caveat: str | None = None + + +OPERATORS: dict[str, Operator] = { + "site": Operator( + "site:", "Site / domain restriction", + "Only return pages whose host matches.", + "The load-bearing operator of every recon dork. Matches on the host " + "portion of the URL as a suffix, so site:example.com also covers " + "www.example.com and dev.example.com. Pair it with a negative " + "-site:www.example.com to surface the subdomains DNS enumeration " + "missed, since Google only indexes hosts it can reach and link to.", + "site:example.com", + ), + "filetype": Operator( + "filetype:", "File type", + "Restrict to one file extension.", + "Matches the extension Google recorded for the document, not the " + "Content-Type header. One extension per operator - stack them with " + "OR to cover a family. Google indexes a fixed list of types (pdf, " + "doc(x), xls(x), ppt(x), rtf, ps, swf, kml, txt and a handful more); " + "a filetype:env dork works only because those files got served as " + "plain text and indexed as such, which is also why it is such a " + "reliable finding.", + "filetype:pdf", + ), + "ext": Operator( + "ext:", "File extension (alias)", + "Synonym for filetype: on Google and Bing.", + "Functionally identical to filetype: on Google, Bing and Brave. " + "Shorter, so it reads better in long stacked dorks. Yandex does not " + "accept it - use mime: there.", + "ext:sql", + ), + "inurl": Operator( + "inurl:", "Term in URL", + "The term appears anywhere in the URL.", + "Substring match against the whole URL including query string, which " + "makes it the operator for finding parameter names worth testing " + "(inurl:id=, inurl:redirect=, inurl:file=). It matches one term; for " + "several terms all of which must appear, use allinurl:.", + "inurl:admin", + ), + "allinurl": Operator( + "allinurl:", "All terms in URL", + "Every following term must appear in the URL.", + "Applies to everything after it, which is why it does not mix with " + "other operators - Google quietly ignores the rest of the query. " + "Prefer stacking several inurl: operators instead; the result is the " + "same and the dork stays composable.", + "allinurl:admin login", + status="flaky", + caveat="Do not combine with other operators - the rest of the query is ignored.", + ), + "intitle": Operator( + "intitle:", "Term in title", + "The term appears in the page <title>.", + "The title is the highest-signal field for fingerprinting an " + "application, because vendors ship a default. intitle:\"index of\" is " + "the classic directory-listing dork; intitle:\"phpMyAdmin\" finds " + "exposed database consoles by their stock title.", + 'intitle:"index of"', + ), + "allintitle": Operator( + "allintitle:", "All terms in title", + "Every following term must appear in the title.", + "Same swallowing behaviour as allinurl: - it consumes the rest of the " + "query and disables other operators. Stack intitle: instead.", + "allintitle:index of backup", + status="flaky", + caveat="Do not combine with other operators.", + ), + "intext": Operator( + "intext:", "Term in body text", + "The term appears in the page body.", + "Forces the term into the body rather than letting Google match it " + "in the title, URL or inbound anchor text. Useful for pinning a " + "string that would otherwise match the navigation of every page on a " + "site, and for secret patterns such as intext:\"BEGIN RSA PRIVATE KEY\".", + 'intext:"password"', + ), + "allintext": Operator( + "allintext:", "All terms in body", + "Every following term must appear in the body.", + "Consumes the rest of the query like the other allin* operators.", + "allintext:username password", + status="flaky", + caveat="Do not combine with other operators.", + ), + "inanchor": Operator( + "inanchor:", "Term in inbound anchor text", + "Matches the text other pages use to link here.", + "Searches the anchor text of inbound links rather than the page " + "itself. Thin and unreliable now that Google has de-emphasised link " + "signals, but occasionally surfaces a page whose own content never " + "names what it is.", + "inanchor:login", + status="flaky", + ), + "phrase": Operator( + '"..."', "Exact phrase", + "Match the quoted string verbatim, in order.", + "Disables stemming, synonym expansion and word reordering. Any dork " + "that hunts for a literal artefact - a banner, an error string, a key " + "header - belongs in quotes, or Google will helpfully find you " + "something adjacent instead of the thing.", + '"index of /backup"', + ), + "or": Operator( + "OR", "Disjunction", + "Match either side. Must be uppercase.", + "Written OR or | . Lowercase 'or' is treated as a stop word and " + "silently ignored, which is the single most common reason a dork " + "returns the wrong thing. Group alternatives in parentheses when you " + "mix them with AND-ed terms: site:x.com (ext:sql | ext:bak).", + "ext:sql OR ext:bak", + ), + "exclude": Operator( + "-term", "Exclusion", + "Drop results containing the term.", + "Prefix any term or operator with a hyphen and no space. The " + "workhorse for cutting noise: -site:www to leave only subdomains, " + "-inurl:wp-content to drop theme assets, -\"404\" to lose soft error " + "pages.", + "-site:www.example.com", + ), + "wildcard": Operator( + "*", "Single-token wildcard", + "Stands in for one or more whole words.", + "Only meaningful inside a quoted phrase, where it matches whole " + "tokens rather than characters - \"admin * panel\" matches 'admin " + "control panel'. It is not a substring wildcard: site:*.example.com " + "works, but adm*n does not.", + '"confidential * report"', + ), + "around": Operator( + "AROUND(n)", "Proximity", + "Two terms within n words of each other.", + "Uppercase, no space before the bracket. Tighter than an AND and " + "looser than a phrase, which makes it the right tool when you know " + "two things co-occur but not in what order - password AROUND(3) " + "database finds the sentence without guessing its shape.", + "password AROUND(5) database", + ), + "before": Operator( + "before:", "Upper date bound", + "Only documents Google dates before YYYY-MM-DD.", + "Filters on Google's estimate of the document date, which is often " + "wrong for pages without explicit markup. Good for narrowing a noisy " + "result set, never good enough to prove when something was published.", + "before:2020-01-01", + ), + "after": Operator( + "after:", "Lower date bound", + "Only documents Google dates after YYYY-MM-DD.", + "Same caveat as before:. Pair the two to bracket a window. The most " + "useful recon framing is after: a known breach or migration date, to " + "see what got published since.", + "after:2024-01-01", + ), + "numrange": Operator( + "n..m", "Numeric range", + "Match any number in the range.", + "Two dots between two integers. Works in plain terms and inside " + "other operators, so inurl:id=1..100 is legal. Legacy numrange:1-100 " + "syntax still parses but the dotted form is the documented one.", + "1000..2000", + ), + "related": Operator( + "related:", "Similar sites", + "Sites Google considers similar.", + "Was a genuine competitor-discovery tool; now returns nothing for " + "most inputs and is officially unsupported in Search. Occasionally " + "still answers for very large domains. Treat any result as a bonus.", + "related:example.com", + status="flaky", + caveat="Officially unsupported since 2023; usually returns nothing.", + ), + "cache": Operator( + "cache:", "Cached copy", + "RETIRED - removed by Google in 2024.", + "Google removed the cache: operator and the cached-page links in " + "September 2024. For a historical copy use the Wayback Machine " + "(web.archive.org/web/*/example.com/*), archive.today, or Bing's own " + "cached link where it still appears.", + "cache:example.com", + status="retired", + caveat="Use web.archive.org or archive.today instead.", + ), + "link": Operator( + "link:", "Inbound links", + "RETIRED - dropped in 2017.", + "Returned pages linking to a URL. Deprecated in 2017 and now a no-op " + "that silently degrades to a plain keyword search. Backlink data " + "lives in commercial SEO tools or Search Console for sites you own.", + "link:example.com", + status="retired", + caveat="No-op. Use Search Console or a backlink tool.", + ), + "info": Operator( + "info:", "Page info", + "RETIRED - folded into normal search.", + "Used to show Google's summary card for a URL. Retired; a bare URL " + "search gets you the same thing.", + "info:example.com", + status="retired", + ), + "define": Operator( + "define:", "Definition", + "Dictionary entry for a term.", + "Not a recon operator, but worth knowing it exists so you do not " + "confuse it with a filter - it short-circuits the whole query into a " + "dictionary card.", + "define:kerberoasting", + ), + "source": Operator( + "source:", "News source", + "Restrict Google News to one publication.", + "Only meaningful on news.google.com. Occasionally useful for " + "mapping an org's public announcements, acquisitions and outage " + "history during a passive-recon pass.", + "source:reuters", + ), + "imagesize": Operator( + "imagesize:", "Exact image dimensions", + "Images at exactly WxH pixels.", + "Images vertical only. Niche, but it is the fastest way to pull " + "every copy of a specific badge, logo or screenshot from an org.", + "imagesize:1920x1080", + ), + # --- non-Google natives, reachable via translation ------------------- + "ip": Operator( + "ip:", "Hosts on an IP (Bing)", + "Bing-only: every indexed site on an address.", + "Bing's single best recon operator and it has no Google equivalent. " + "ip:203.0.113.10 lists the virtual hosts Bing saw on that address - a " + "free reverse-IP lookup that often reveals the rest of a shared " + "estate behind one origin.", + "ip:203.0.113.10", + ), + "contains": Operator( + "contains:", "Links to a file type (Bing)", + "Bing-only: pages linking to that extension.", + "Different from filetype:. contains:sql returns pages that *link to* " + "a .sql file rather than the file itself, which catches the index " + "page of a backup directory that was never itself indexed.", + "contains:sql", + ), + "host": Operator( + "host:", "Exact host (Yandex)", + "Yandex-only: exact host match.", + "Yandex distinguishes host: (this exact host) from rhost: (reversed " + "domain, so rhost:com.example.* covers the whole tree). Both are " + "stricter and more predictable than Google's suffix matching.", + "host:dev.example.com", + ), + "mime": Operator( + "mime:", "MIME type (Yandex)", + "Yandex-only: filter by document type.", + "Yandex's answer to filetype:, keyed on the recorded MIME type.", + "mime:pdf", + ), + "org": Operator( + "org:", "GitHub organisation", + "Code search scoped to one org.", + "GitHub code search. Combine with path:, language: and a literal " + "string to hunt committed secrets in an org's public repositories. " + "Requires being signed in.", + "org:example-inc", + ), + "path": Operator( + "path:", "GitHub file path", + "Code search scoped to a path or filename.", + "Replaced the old filename: operator. path:.env matches any file at " + "that path segment; path:**/config/*.yml globs.", + "path:.env", + ), + "language": Operator( + "language:", "GitHub language", + "Code search scoped to one language.", + "Narrows a noisy content search to the language the secret would " + "plausibly live in.", + "language:python", + ), +} + +RETIRED = {k for k, v in OPERATORS.items() if v.status == "retired"} + + +# --------------------------------------------------------------------------- +# Engines +# --------------------------------------------------------------------------- +@dataclass(frozen=True) +class Engine: + key: str + label: str + url: str # printf-style with {q} + supports: frozenset[str] + rewrites: dict[str, str] = field(default_factory=dict) + notes: str = "" + login_warning: bool = False + + +_CORE = {"site", "filetype", "ext", "inurl", "intitle", "intext", "phrase", + "or", "exclude", "wildcard"} + +ENGINES: dict[str, Engine] = { + "google": Engine( + "google", "Google", + "https://www.google.com/search?q={q}", + frozenset(_CORE | {"allinurl", "allintitle", "allintext", "inanchor", + "around", "before", "after", "numrange", "related", + "define", "source", "imagesize"}), + notes="The reference dialect. Everything in this tool is written in it.", + login_warning=True, + ), + "bing": Engine( + "bing", "Bing", + "https://www.bing.com/search?q={q}", + frozenset(_CORE | {"ip", "contains"}), + rewrites={"intext:": "inbody:"}, + notes="Adds ip: (reverse-IP) and contains: (links to a filetype). " + "intext: becomes inbody:. No AROUND(), no before:/after:.", + ), + "ddg": Engine( + "ddg", "DuckDuckGo", + "https://duckduckgo.com/?q={q}", + frozenset(_CORE - {"wildcard"}), + notes="Relays Bing's index with its own ranking. Operator support is " + "real but shallow: no AROUND(), no date bounds, wildcards are " + "inert. Its bangs (!gh, !so) are shortcuts, not operators.", + ), + "yandex": Engine( + "yandex", "Yandex", + "https://yandex.com/search/?text={q}", + frozenset({"site", "phrase", "or", "exclude", "host", "mime"}), + rewrites={"filetype:": "mime:", "ext:": "mime:", "intitle:": "title:"}, + notes="Often indexes hosts Google never crawled, which makes it the " + "second engine worth running any subdomain dork through. Its " + "native spellings differ: title: not intitle:, mime: not " + "filetype:, url: not inurl:, plus host:/rhost:.", + ), + "brave": Engine( + "brave", "Brave Search", + "https://search.brave.com/search?q={q}", + frozenset(_CORE), + notes="Independent index, no account required, no personalisation - " + "the lowest-friction engine to dork from a clean browser.", + ), + "github": Engine( + "github", "GitHub code search", + "https://github.com/search?type=code&q={q}", + frozenset({"org", "path", "language", "phrase", "or", "exclude"}), + notes="A different dialect entirely: org:, repo:, path:, language:, " + "symbol:. Requires a signed-in account - use a throwaway, never " + "your real one.", + login_warning=True, + ), + "shodan": Engine( + "shodan", "Shodan", + "https://www.shodan.io/search?query={q}", + frozenset({"phrase", "exclude"}), + notes="Indexes service banners, not web pages. Its own keys " + "(ssl.cert.subject.cn:, http.title:, http.html:, org:, net:, " + "port:) replace the web operators outright.", + login_warning=True, + ), +} + + +# --------------------------------------------------------------------------- +# The dork library +# --------------------------------------------------------------------------- +@dataclass(frozen=True) +class Recipe: + template: str + why: str + noise: str = "low" # low | medium | high - how junky the results are + + +@dataclass(frozen=True) +class Objective: + key: str + label: str + glyph: str + blurb: str + needs: tuple[str, ...] # ordered subset of domain/org/keyword/ext + recipes: tuple[Recipe, ...] + caution: str | None = None + + +OBJECTIVES: tuple[Objective, ...] = ( + Objective( + "files", "Exposed files & configs", "file", + "Config, environment, log and key files that were never meant to be " + "served - the highest-yield category in the whole practice.", + ("domain",), + ( + Recipe( + 'site:{domain} ext:env OR ext:cfg OR ext:conf OR ext:ini', + "Application config. A served .env is credentials in plaintext - " + "database URI, mail password, cloud keys, app secret.", + ), + Recipe( + 'site:{domain} ext:log OR ext:txt inurl:log', + "Application and error logs. Usually leak internal paths, " + "usernames, session tokens and stack traces.", + noise="medium", + ), + Recipe( + 'site:{domain} ext:bak OR ext:old OR ext:backup OR ext:swp OR ext:save', + "Editor and deploy leftovers. index.php.bak is served as text " + "instead of executed, which hands over the source.", + ), + Recipe( + 'site:{domain} ext:yml OR ext:yaml OR ext:toml OR ext:json inurl:config', + "Modern config formats - docker-compose, CI definitions, " + "appsettings.json, secrets baked into a values file.", + noise="medium", + ), + Recipe( + 'site:{domain} ext:pem OR ext:key OR ext:ppk OR ext:p12 OR ext:pfx', + "Private keys and certificate bundles. Any hit here is a " + "critical finding on its own.", + ), + Recipe( + 'site:{domain} inurl:.git OR inurl:.svn OR inurl:.hg', + "Exposed VCS metadata. A readable .git/ directory means the " + "whole repository and its history can be reconstructed.", + ), + Recipe( + 'site:{domain} inurl:wp-config OR inurl:configuration.php OR inurl:settings.py', + "Framework config by its canonical filename - WordPress, " + "Joomla, Django.", + ), + Recipe( + 'site:{domain} ext:sql OR ext:dbf OR ext:mdb OR ext:sqlite', + "Database files served straight off the webroot.", + ), + ), + caution="A file being indexed does not make it yours. Read the result " + "snippet; do not download what is out of scope.", + ), + Objective( + "panels", "Login & admin panels", "lock", + "Authentication surfaces, management consoles and anything that " + "should have been behind the VPN.", + ("domain",), + ( + Recipe( + 'site:{domain} inurl:admin OR inurl:administrator OR inurl:adminpanel', + "The obvious path names. Still hits constantly.", + noise="medium", + ), + Recipe( + 'site:{domain} intitle:"login" OR intitle:"sign in" OR intitle:"log in"', + "Title-based, so it catches panels on paths you would never " + "guess.", + noise="medium", + ), + Recipe( + 'site:{domain} inurl:phpmyadmin OR inurl:adminer OR inurl:pma', + "Exposed database consoles. Frequently unauthenticated or on " + "vendor defaults.", + ), + Recipe( + 'site:{domain} intitle:"Dashboard" (inurl:jenkins OR inurl:grafana OR inurl:kibana)', + "CI and observability dashboards - build secrets, internal " + "hostnames, query access to production telemetry.", + ), + Recipe( + 'site:{domain} inurl:/manager/html OR inurl:/axis2 OR inurl:/jmx-console', + "Java application-server managers. Tomcat's manager with " + "default creds is a WAR deploy away from RCE.", + ), + Recipe( + 'site:{domain} inurl:owa OR inurl:/rdweb OR inurl:citrix OR inurl:vpn', + "Remote-access front doors - the surface password spraying " + "actually targets.", + ), + Recipe( + 'site:{domain} intitle:"index of" inurl:admin', + "A browsable admin directory: the panel and its source at " + "once.", + ), + ), + caution="Finding a panel is recon. Logging in - even with default " + "credentials - is access, and needs it in writing.", + ), + Objective( + "listing", "Open directory listings", "folder", + "Webservers with autoindex on, handing out a file browser.", + ("domain",), + ( + Recipe( + 'site:{domain} intitle:"index of"', + "The canonical dork. Apache, nginx and IIS all emit a title " + "that starts this way.", + noise="medium", + ), + Recipe( + 'site:{domain} intitle:"index of" (backup OR bak OR old OR archive)', + "Listings that name themselves as a dumping ground.", + ), + Recipe( + 'site:{domain} intitle:"index of" "parent directory" (.sql OR .zip OR .tar.gz)', + "Listings with an archive or dump sitting in them.", + ), + Recipe( + 'site:{domain} intitle:"index of" (uploads OR files OR documents)', + "User-upload directories - often unfiltered, sometimes " + "writable.", + noise="medium", + ), + Recipe( + 'site:{domain} "Directory Listing For" OR "Index of /" -inurl:html', + "Tomcat and other servers whose listing wording differs from " + "Apache's.", + ), + ), + ), + Objective( + "cloud", "Cloud storage buckets", "cloud", + "Public object storage referenced by, or belonging to, the target.", + ("domain", "org"), + ( + Recipe( + 'site:s3.amazonaws.com "{org}"', + "S3 buckets whose listing mentions the organisation.", + ), + Recipe( + 'site:{domain} inurl:s3.amazonaws.com OR inurl:s3-external', + "Pages on the target that link into S3 - the fastest way to " + "learn the bucket naming convention.", + ), + Recipe( + 'inurl:blob.core.windows.net "{org}"', + "Azure Blob containers.", + ), + Recipe( + 'inurl:storage.googleapis.com "{org}"', + "Google Cloud Storage objects.", + ), + Recipe( + 'site:{domain} inurl:digitaloceanspaces.com OR inurl:r2.dev OR inurl:backblazeb2.com', + "The smaller providers, which get less scrutiny and are " + "misconfigured more often.", + ), + Recipe( + '"{org}" (site:trello.com OR site:notion.site OR site:docs.google.com)', + "Public boards and documents. Trello boards in particular " + "leak credentials and architecture diagrams at a remarkable " + "rate.", + noise="medium", + ), + ), + caution="Listing a public bucket is passive. Enumerating or " + "downloading its contents is not, and may be unauthorised " + "access even when the bucket is world-readable.", + ), + Objective( + "secrets", "Credentials, keys & tokens", "key", + "Live secret material indexed in pages, files or code.", + ("domain", "org"), + ( + Recipe( + 'site:{domain} intext:"BEGIN RSA PRIVATE KEY" OR intext:"BEGIN OPENSSH PRIVATE KEY"', + "Private key headers in page text - pasted into a wiki, a " + "ticket, a README.", + ), + Recipe( + 'site:{domain} intext:"api_key" OR intext:"apikey" OR intext:"client_secret"', + "Generic secret variable names in served content.", + noise="high", + ), + Recipe( + 'site:{domain} (ext:env OR ext:yml) intext:PASSWORD', + "Config files with a password assignment in them.", + ), + Recipe( + 'site:{domain} intext:"AKIA" ', + "AWS access key IDs all begin AKIA - a distinctive literal " + "that rarely false-positives.", + ), + Recipe( + '"{org}" ("xoxb-" OR "xoxp-" OR "ghp_" OR "sk_live_")', + "Slack, GitHub and Stripe token prefixes. Each is " + "structurally unique, so a hit is almost never noise.", + ), + Recipe( + 'site:pastebin.com OR site:ghostbin.co OR site:justpaste.it "{domain}"', + "Paste sites - where dumps, configs and credential lists get " + "parked.", + noise="medium", + ), + Recipe( + 'site:{domain} inurl:/.aws/credentials OR inurl:/.ssh/id_rsa OR inurl:/.npmrc', + "Dotfiles that made it into a webroot.", + ), + ), + caution="A live key is a reportable finding, not a login. Do not " + "authenticate with anything you find; report it immediately, " + "and treat the search itself as sensitive.", + ), + Objective( + "docs", "Documents & metadata", "book", + "Published files whose contents - or EXIF/author metadata - give up " + "names, software versions and internal paths.", + ("domain", "org"), + ( + Recipe( + 'site:{domain} ext:pdf OR ext:docx OR ext:xlsx OR ext:pptx', + "The corpus. Pull it, then run exiftool over the lot for " + "Author, Creator and internal UNC paths.", + noise="medium", + ), + Recipe( + 'site:{domain} ext:pdf (confidential OR internal OR "not for distribution")', + "Documents labelled as restricted that are nonetheless " + "public.", + ), + Recipe( + 'site:{domain} ext:xlsx (employee OR salary OR roster OR contact)', + "Spreadsheets - the format people paste staff lists into.", + ), + Recipe( + 'site:{domain} ext:vsd OR ext:vsdx OR ext:drawio OR ext:mmd', + "Network and architecture diagrams.", + ), + Recipe( + '"{org}" filetype:pdf intext:"@{domain}"', + "Documents from anywhere that contain the target's email " + "addresses - harvests the address format.", + noise="medium", + ), + ), + ), + Objective( + "errors", "Error messages & stack traces", "bug", + "Verbose failures: the cheapest source of stack depth, framework " + "versions and absolute paths.", + ("domain",), + ( + Recipe( + 'site:{domain} intext:"Warning: mysql_connect()" OR intext:"You have an error in your SQL syntax"', + "Database errors. The second string is the literal MySQL " + "parser message that confirms injectable input.", + ), + Recipe( + 'site:{domain} intext:"Fatal error" OR intext:"Uncaught exception"', + "PHP and Java fatals - usually carry the absolute filesystem " + "path.", + noise="medium", + ), + Recipe( + 'site:{domain} intext:"Traceback (most recent call last)"', + "Python tracebacks. Django debug pages additionally dump " + "settings and environment.", + ), + Recipe( + 'site:{domain} intitle:"Whoops, looks like something went wrong"', + "Laravel's debug page, which shows environment variables " + "verbatim.", + ), + Recipe( + 'site:{domain} intext:"Server Error in" "Application" intext:"Stack Trace"', + "ASP.NET yellow-screen-of-death with the trace enabled.", + ), + Recipe( + 'site:{domain} intext:"phpinfo()" OR intitle:"phpinfo()"', + "A left-behind phpinfo page - full build config, loaded " + "modules, absolute paths, sometimes environment secrets.", + ), + ), + ), + Objective( + "db", "Databases, backups & dumps", "db", + "Whole datasets, in whatever format they escaped as.", + ("domain", "org"), + ( + Recipe( + 'site:{domain} ext:sql intext:"INSERT INTO"', + "SQL dumps, confirmed by their own statement syntax rather " + "than just an extension.", + ), + Recipe( + 'site:{domain} ext:sql intext:"CREATE TABLE" intext:password', + "Dumps that include a credentials table.", + ), + Recipe( + 'site:{domain} (ext:zip OR ext:tar OR ext:gz OR ext:7z) (backup OR dump OR export)', + "Archived backups sitting in the webroot.", + noise="medium", + ), + Recipe( + 'site:{domain} inurl:backup OR inurl:dump OR inurl:export', + "Path-based, catches the directory even when the file itself " + "was not indexed.", + noise="medium", + ), + Recipe( + '"{org}" (site:pastebin.com OR site:anonfiles.com) (dump OR leak OR combo)', + "Third-party hosting of a leak that names the org.", + noise="medium", + ), + ), + caution="Downloading a customer database is not passive recon under " + "any framework. Note the URL, report it, stop.", + ), + Objective( + "params", "Injectable parameters", "bolt", + "URLs carrying the parameter shapes that map to a vulnerability " + "class - a target list for later, authorised testing.", + ("domain",), + ( + Recipe( + 'site:{domain} inurl:id= OR inurl:pid= OR inurl:cat= OR inurl:item=', + "Numeric identifiers - the classic SQL injection and IDOR " + "surface.", + noise="medium", + ), + Recipe( + 'site:{domain} inurl:file= OR inurl:page= OR inurl:path= OR inurl:template=', + "Path-ish parameters - local file inclusion and traversal.", + noise="medium", + ), + Recipe( + 'site:{domain} inurl:url= OR inurl:redirect= OR inurl:next= OR inurl:return=', + "Destination parameters - open redirect, and the front half " + "of many SSRF chains.", + noise="medium", + ), + Recipe( + 'site:{domain} inurl:cmd= OR inurl:exec= OR inurl:query= OR inurl:search=', + "Command and query parameters - injection and reflected XSS.", + noise="medium", + ), + Recipe( + 'site:{domain} inurl:debug=true OR inurl:test= OR inurl:admin=1', + "Flags that flip an application into a more verbose or more " + "privileged mode.", + ), + Recipe( + 'site:{domain} ext:php inurl:? ', + "Every indexed PHP endpoint that takes a query string - the " + "raw list to feed a parameter miner.", + noise="high", + ), + ), + caution="This builds a target list, nothing more. Testing any of it " + "requires the engagement to cover it.", + ), + Objective( + "devices", "Cameras, printers & ICS", "camera", + "Embedded web interfaces indexed by their stock titles and banners.", + ("domain", "keyword"), + ( + Recipe( + 'intitle:"webcamXP" OR intitle:"Live View / - AXIS" OR inurl:/view.shtml', + "IP camera interfaces by vendor default title.", + ), + Recipe( + 'intitle:"HP LaserJet" OR intitle:"Brother" inurl:printer', + "Networked printers - address books, stored jobs, sometimes " + "LDAP credentials in the config page.", + ), + Recipe( + 'intitle:"index of" inurl:ftp "{keyword}"', + "Anonymous FTP roots reachable over HTTP.", + ), + Recipe( + 'inurl:/level/15/exec OR intitle:"Cisco" inurl:/cgi-bin', + "Network-device management CGIs.", + ), + Recipe( + 'site:{domain} intitle:"Router" OR intitle:"Gateway" inurl:status', + "Edge devices on the target's own domain.", + ), + ), + caution="Most device dorks are untargeted by nature - they return " + "hardware belonging to strangers. Scope them with site: or " + "do not run them. Never interact with a device you do not " + "have authorisation for.", + ), + Objective( + "people", "Employees, emails & usernames", "users", + "The identity layer: names, address format, and the username " + "convention that feeds spraying lists.", + ("domain", "org"), + ( + Recipe( + 'site:linkedin.com/in "{org}"', + "Staff profiles. Names plus role, which is what you need to " + "derive the username scheme.", + noise="medium", + ), + Recipe( + 'site:{domain} intext:"@{domain}"', + "The organisation's own pages leaking its address format " + "(first.last@, finitial.last@, flast@).", + noise="medium", + ), + Recipe( + '"{org}" (site:github.com OR site:gitlab.com) intext:"@{domain}"', + "Developers using a work address in commits or profiles.", + ), + Recipe( + '"{org}" (site:stackoverflow.com OR site:serverfault.com)', + "Engineers asking about their own stack - frequently pasting " + "real config with real hostnames.", + noise="medium", + ), + Recipe( + '"{org}" ("we use" OR "powered by" OR "built with") -site:{domain}', + "Third-party mentions of the tech stack: job ads, case " + "studies, conference talks.", + noise="high", + ), + Recipe( + 'site:{domain} (inurl:team OR inurl:staff OR inurl:about) intext:"@"', + "The org's own directory page.", + ), + ), + caution="Personal data is in scope for GDPR and equivalents whether " + "or not it is in scope for the test. Collect the minimum, " + "store it with the engagement, delete it after.", + ), + Objective( + "code", "Code, repos & CI", "code", + "Source and pipeline configuration, mostly through GitHub's own " + "dialect.", + ("domain", "org"), + ( + Recipe( + 'org:{org} path:.env', + "Committed environment files, scoped to the organisation.", + ), + Recipe( + 'org:{org} "BEGIN RSA PRIVATE KEY"', + "Keys in the org's public code.", + ), + Recipe( + '"{domain}" "password" language:yaml', + "Credentials in YAML anywhere on GitHub that reference the " + "target's domain.", + noise="medium", + ), + Recipe( + 'org:{org} path:.github/workflows', + "CI definitions - runner labels, deploy targets, secret " + "names, and occasionally the secrets themselves.", + ), + Recipe( + '"{domain}" ("192.168." OR "10.0." OR ".local")', + "Internal addressing leaked into public code - a free " + "network map.", + noise="medium", + ), + Recipe( + 'org:{org} path:Dockerfile', + "Base images, build args and package versions.", + ), + ), + caution="Public repositories only. Cloning a private repo you can " + "somehow reach is unauthorised access, full stop.", + ), + Objective( + "subdomains", "Subdomains & forgotten hosts", "globe", + "Hosts the index knows about that DNS enumeration did not return.", + ("domain",), + ( + Recipe( + 'site:*.{domain} -site:www.{domain}', + "The core subdomain dork: everything under the apex except " + "the main site.", + ), + Recipe( + 'site:*.{domain} (dev OR staging OR test OR uat OR qa OR beta)', + "Non-production environments - weaker credentials, debug on, " + "real data.", + ), + Recipe( + 'site:*.{domain} (vpn OR remote OR portal OR mail OR intranet)', + "Access infrastructure and internal-facing hosts.", + ), + Recipe( + 'site:*.{domain} -site:www.{domain} -site:blog.{domain} -site:shop.{domain}', + "Iteratively subtract the hosts you already know to force " + "new ones to the surface.", + ), + Recipe( + '"{domain}" -site:{domain}', + "Third-party pages that reference the domain - partners, " + "status pages, monitoring, archived copies.", + noise="high", + ), + ), + ), + Objective( + "archive", "Archived & removed pages", "book", + "Content that was taken down but still exists somewhere.", + ("domain", "keyword"), + ( + Recipe( + 'site:web.archive.org "{domain}"', + "Wayback captures indexed by Google. For the full history " + "query the CDX API directly instead.", + noise="medium", + ), + Recipe( + 'site:{domain} before:2022-01-01 "{keyword}"', + "Older material on the live site, bounded by Google's date " + "estimate.", + ), + Recipe( + 'site:archive.today "{domain}" OR site:archive.ph "{domain}"', + "The other archive, which captures pages Wayback refuses.", + ), + Recipe( + 'cache:{domain}', + "RETIRED - kept here so the tool can show you what replaced " + "it. Use the Wayback Machine.", + ), + ), + ), +) + +OBJECTIVE_BY_KEY = {o.key: o for o in OBJECTIVES} + +FIELD_PROMPTS = { + "domain": ("Target domain", "example.com", "apex domain, no scheme, no path"), + "org": ("Organisation name", "Example Inc", "as it appears in profiles and docs"), + "keyword": ("Keyword", "invoice", "a term specific to this hunt"), + "ext": ("File extension", "pdf", "without the dot"), +} + + +# --------------------------------------------------------------------------- +# Query analysis +# --------------------------------------------------------------------------- +_OP_TOKEN = re.compile(r"(?<![\w.])-?([a-zA-Z_]+):") +_ALIASES = {"inbody": "intext", "rhost": "host", "filename": "path"} + + +def used_operators(query: str) -> list[str]: + """Operator keys present in a query, in order of first appearance.""" + seen: list[str] = [] + + def add(key: str) -> None: + if key in OPERATORS and key not in seen: + seen.append(key) + + for m in _OP_TOKEN.finditer(query): + name = m.group(1).lower() + add(_ALIASES.get(name, name)) + if '"' in query: + add("phrase") + if re.search(r"\bOR\b|\|", query): + add("or") + if re.search(r"(?:^|\s)-\S", query): + add("exclude") + if "*" in query: + add("wildcard") + if re.search(r"AROUND\(\d+\)", query): + add("around") + if re.search(r"\d+\.\.\d+", query): + add("numrange") + return seen + + +# Where an operator has no home on an engine, say what to do instead. +_FALLBACK: dict[tuple[str, str], str] = { + ("yandex", "intext"): "drop it - Yandex searches the body by default", + ("yandex", "inurl"): "not Yandex's spelling - use url:example.com/* instead", + ("ddg", "wildcard"): "inert here; tighten with a quoted phrase", + ("github", "site"): "use org: or repo: instead", + ("github", "filetype"): "use path: instead", + ("github", "ext"): "use path: instead", + ("shodan", "site"): "use hostname: or ssl.cert.subject.cn: instead", + ("shodan", "intitle"): "use http.title: instead", + ("shodan", "intext"): "use http.html: instead", +} +_GENERIC_FALLBACK = { + "around": "no proximity operator - use a quoted phrase", + "before": "no date operator - use the engine's date filter UI", + "after": "no date operator - use the engine's date filter UI", + "numrange": "no range syntax - enumerate the values", + "inanchor": "no anchor-text operator", + "related": "no similarity operator", +} + + +def translate(query: str, engine: Engine) -> tuple[str, list[tuple[str, str]]]: + """Rewrite a Google-dialect query for `engine`. + + Returns the rewritten query and a list of (operator, note) pairs for + everything that did not survive the trip intact. + """ + out = query + notes: list[tuple[str, str]] = [] + + for src, dst in engine.rewrites.items(): + pattern = re.compile(re.escape(src), re.IGNORECASE) + if pattern.search(out): + out = pattern.sub(dst, out) + notes.append((src, f"rewritten to {dst} for {engine.label}")) + + for key in used_operators(out): + if key in engine.supports: + continue + op = OPERATORS[key] + advice = _FALLBACK.get((engine.key, key)) or _GENERIC_FALLBACK.get(key) + if advice is None: + advice = f"not supported on {engine.label} - the terms still search, unfiltered" + notes.append((op.key, advice)) + return out, notes + + +def highlight(query: str) -> Text: + t = Text(query, style="df.text") + for m in re.finditer(r'"[^"]*"', query): + t.stylize("df.val", m.start(), m.end()) + for m in _OP_TOKEN.finditer(query): + t.stylize("df.op", m.start(), m.end()) + for m in re.finditer(r"\bOR\b|\|", query): + t.stylize("df.kw", m.start(), m.end()) + for m in re.finditer(r"AROUND\(\d+\)", query): + t.stylize("df.kw", m.start(), m.end()) + for m in re.finditer(r"(?:^|\s)(-)[\w\"]", query): + t.stylize("df.love", m.start(1), m.end(1)) + for m in re.finditer(r"\*", query): + t.stylize("df.rose", m.start(), m.end()) + return t + + +def search_url(engine: Engine, query: str) -> str: + return engine.url.format(q=urllib.parse.quote_plus(query)) + + +# --------------------------------------------------------------------------- +# Host integration: clipboard, browser, session log, state +# --------------------------------------------------------------------------- +_CLIPBOARD = ( + ("pbcopy", ["pbcopy"]), + ("wl-copy", ["wl-copy"]), + ("xclip", ["xclip", "-selection", "clipboard"]), + ("xsel", ["xsel", "--clipboard", "--input"]), + ("clip.exe", ["clip.exe"]), +) + + +def clipboard_tool() -> tuple[str, list[str]] | None: + for name, cmd in _CLIPBOARD: + if shutil.which(cmd[0]): + return name, cmd + return None + + +def copy_to_clipboard(text: str) -> tuple[bool, str]: + found = clipboard_tool() + if found is None: + return False, "no clipboard helper found (pbcopy/wl-copy/xclip/xsel/clip.exe)" + name, cmd = found + try: + subprocess.run(cmd, input=text.encode(), check=True) + except (OSError, subprocess.CalledProcessError) as exc: + return False, f"{name} failed: {exc}" + return True, name + + +def state_dir() -> Path: + root = os.environ.get("XDG_STATE_HOME") + base = Path(root) if root else Path.home() / ".local" / "state" + return base / "dorkforge" + + +def ack_path() -> Path: + return state_dir() / "ack.json" + + +def load_ack() -> dict | None: + try: + return json.loads(ack_path().read_text(encoding="utf-8")) + except (OSError, ValueError): + return None + + +def save_ack(scope: str) -> None: + path = ack_path() + try: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text( + json.dumps( + { + "acknowledged": True, + "at": datetime.now(timezone.utc).isoformat(timespec="seconds"), + "scope_note": scope, + "version": VERSION, + }, + indent=2, + ), + encoding="utf-8", + ) + except OSError: + pass # a read-only home should not stop the tool working + + +class SessionLog: + """Append-only markdown log of everything built this run.""" + + def __init__(self, path: Path, scope: str): + self.path = path + self.scope = scope + self.count = 0 + self._started = path.exists() + + def _header(self) -> str: + return ( + f"# dorkforge session\n\n" + f"- **Started:** {datetime.now().astimezone():%Y-%m-%d %H:%M %Z}\n" + f"- **Authorised scope:** {self.scope or 'not recorded'}\n" + f"- **Tool:** dorkforge {VERSION}\n\n" + f"---\n\n" + ) + + def add(self, engine: Engine, query: str, why: str, notes: list[tuple[str, str]]) -> None: + chunk = [] + if not self._started: + chunk.append(self._header()) + self._started = True + self.count += 1 + chunk.append(f"## {self.count}. {engine.label}\n\n") + if why: + chunk.append(f"{why}\n\n") + chunk.append(f"```text\n{query}\n```\n\n") + chunk.append(f"<{search_url(engine, query)}>\n\n") + if notes: + chunk.append("Translation notes:\n\n") + for op, note in notes: + chunk.append(f"- `{op}` - {note}\n") + chunk.append("\n") + chunk.append(f"_Built {datetime.now().astimezone():%Y-%m-%d %H:%M}._\n\n") + with self.path.open("a", encoding="utf-8") as fh: + fh.write("".join(chunk)) + + +# --------------------------------------------------------------------------- +# Rendering +# --------------------------------------------------------------------------- +PLACEHOLDERS = {"domain": "$DOMAIN", "org": "$ORG", "keyword": "$KEYWORD", "ext": "$EXT"} + + +def as_template(text: str) -> str: + return text.format(**PLACEHOLDERS) + + +def banner(console: Console) -> None: + lock = Text() + lock.append(" ", style="df.love") + lock.append(g("dot") + " ", style="df.love") + lock.append("DÆMON", style="df.head") + lock.append("//", style="df.love") + lock.append("SEC", style="df.head") + sub = Text() + sub.append(" dorkforge ", style="df.foam") + sub.append(f"v{VERSION}", style="df.faint") + sub.append(" " + g("search") + " ", style="df.iris") + sub.append("search-operator workbench", style="df.dim") + console.print() + console.print(lock) + console.print(sub) + console.print(Rule(style="df.rule")) + + +def gate(console: Console, force: bool = False) -> str: + """Banner + one-time authorisation acknowledgement. Returns a scope note.""" + prior = load_ack() + if prior and not force: + note = prior.get("scope_note") or "" + line = Text(" " + g("shield") + " authorisation acknowledged ", style="df.foam") + line.append(prior.get("at", "")[:10], style="df.faint") + if note: + line.append(f" {g('dot')} {note}", style="df.faint") + console.print(line) + console.print() + return note + + body = Text() + body.append("dorkforge builds search queries. It sends nothing to any target " + "and needs no API key.\n\n", style="df.dim") + body.append("What it cannot do is make a search legal. ", style="df.text") + body.append( + "Indexed does not mean authorised: retrieving a config file, " + "enumerating a bucket, or logging into a panel you found is access, " + "and access needs written permission. Personal data you collect is " + "regulated whether or not the engagement covers it.\n\n", + style="df.dim", + ) + body.append("Use this against your own estate, a scope you hold in writing, " + "or a lab you are entitled to.", style="df.text") + console.print( + Panel(body, title=Text(f" {g('warn')} BEFORE YOU START ", style="df.warn"), + title_align="left", border_style="df.gold", box=box.SQUARE, padding=(1, 2)) + ) + console.print() + + ok = questionary.confirm( + "I have authorisation for what I am about to look for.", + default=False, style=QUESTIONARY_STYLE, auto_enter=False, + ).ask() + if not ok: + console.print(Text(f" {g('cross')} Not acknowledged - nothing to do.", style="df.bad")) + raise SystemExit(1) + + # No `default=` here: questionary pre-fills it and parks the cursor at the + # end, so anything typed lands *after* the default instead of replacing it. + scope = questionary.text( + "Scope reference - recorded in every session log:", + instruction="(engagement ID, ticket, or leave blank for 'personal lab') ", + style=QUESTIONARY_STYLE, qmark=g("shield"), + ).ask() or "" + scope = scope.strip() or "personal lab" + save_ack(scope) + console.print(Text(f" {g('check')} Acknowledged as '{scope}'. Not asking again " + f"(reset with --reset-ack).", style="df.ok")) + console.print() + return scope + + +def show_dork(console: Console, engine: Engine, query: str, why: str, + notes: list[tuple[str, str]]) -> None: + inner = [highlight(query)] + if why: + inner += [Text(""), Text(why, style="df.dim")] + console.print() + console.print( + Panel( + Group(*inner), + title=Text(f" {g('search')} {engine.label} ", style="df.foam"), + title_align="left", border_style="df.rule", box=box.SQUARE, + padding=(1, 2), + ) + ) + if notes: + console.print(Text(f" {g('warn')} dialect notes", style="df.warn")) + for op, note in notes: + line = Text(" " + g("dot") + " ", style="df.faint") + line.append(op, style="df.op") + line.append(" " + note, style="df.dim") + console.print(line) + + +def explain(console: Console, query: str, engine: Engine) -> None: + keys = used_operators(query) + if not keys: + console.print(Text(" no operators in this query - it is a plain keyword search.", + style="df.faint")) + return + table = Table(box=box.SIMPLE_HEAD, border_style="df.rule", pad_edge=False, + header_style="df.eyebrow", expand=False) + table.add_column("OPERATOR", style="df.op", no_wrap=True) + table.add_column("WHAT IT DOES", style="df.text") + table.add_column(engine.label.upper(), justify="center", no_wrap=True) + for key in keys: + op = OPERATORS[key] + if op.status == "retired": + mark = Text(g("cross"), style="df.bad") + elif key not in engine.supports: + mark = Text(g("cross"), style="df.love") + elif op.status == "flaky": + mark = Text("~", style="df.gold") + else: + mark = Text(g("check"), style="df.ok") + summary = op.summary + if op.caveat: + summary += f"\n{g('warn')} {op.caveat}" + table.add_row(op.key, summary, mark) + console.print() + console.print(table) + legend = Text(" ", style="df.faint") + legend.append(g("check") + " supported ", style="df.ok") + legend.append("~ unreliable ", style="df.gold") + legend.append(g("cross") + " unsupported or retired", style="df.love") + console.print(legend) + + +def operator_reference(console: Console, only: str | None = None) -> None: + if only: + key = only.strip().rstrip(":").lower() + key = _ALIASES.get(key, key) + op = OPERATORS.get(key) + if op is None: + console.print(Text(f" unknown operator '{only}'. Try --operators for the list.", + style="df.bad")) + raise SystemExit(2) + head = Text(" " + op.key + " ", style="df.op") + head.append(op.label, style="df.head") + if op.status != "live": + head.append(f" [{op.status}]", style="df.warn" if op.status == "flaky" else "df.bad") + console.print() + console.print(head) + console.print(Rule(style="df.rule")) + for line in textwrap.wrap(op.detail, width=min(88, console.width - 4)): + console.print(Text(" " + line, style="df.dim")) + console.print() + ex = Text(" example ", style="df.eyebrow") + ex.append_text(highlight(op.example)) + console.print(ex) + support = [e.label for e in ENGINES.values() if key in e.supports] + console.print(Text(" engines " + (", ".join(support) or "none"), style="df.faint")) + console.print() + return + + table = Table(box=box.SIMPLE_HEAD, border_style="df.rule", header_style="df.eyebrow", + pad_edge=False) + table.add_column("OPERATOR", style="df.op", no_wrap=True) + table.add_column("DOES", style="df.text") + table.add_column("EXAMPLE", style="df.val") + table.add_column("STATUS", no_wrap=True) + for op in OPERATORS.values(): + status = { + "live": Text("live", style="df.ok"), + "flaky": Text("unreliable", style="df.gold"), + "retired": Text("retired", style="df.bad"), + }[op.status] + table.add_row(op.key, op.summary, op.example, status) + console.print() + console.print(table) + console.print(Text(f"\n {g('info')} dorkforge --explain site: for the long form on any one " + f"of these.\n", style="df.faint")) + + +def library_dump(console: Console) -> None: + console.print() + for obj in OBJECTIVES: + head = Text(" " + g(obj.glyph) + " ", style="df.foam") + head.append(obj.label, style="df.head") + head.append(f" ({obj.key})", style="df.faint") + console.print(head) + console.print(Text(" " + obj.blurb, style="df.faint")) + console.print() + for r in obj.recipes: + line = Text(" ") + line.append_text(highlight(as_template(r.template))) + console.print(line) + for wrapped in textwrap.wrap(r.why, width=min(84, console.width - 8)): + console.print(Text(" " + wrapped, style="df.dim")) + console.print() + if obj.caution: + console.print(Text(" " + g("warn") + " " + obj.caution, style="df.warn")) + console.print() + console.print(Rule(style="df.rule")) + console.print(Text(f" {len(OBJECTIVES)} objectives, " + f"{sum(len(o.recipes) for o in OBJECTIVES)} recipes, " + f"{len(OPERATORS)} operators.\n", style="df.faint")) + + +def doctor(console: Console) -> int: + banner(console) + # ok=None means informational: reported, but never a failing check. + rows: list[tuple[str, bool | None, str]] = [] + clip = clipboard_tool() + rows.append(("clipboard", clip is not None, + clip[0] if clip else "install pbcopy / wl-copy / xclip / xsel")) + try: + browser = webbrowser.get() + rows.append(("browser", True, getattr(browser, "name", type(browser).__name__))) + except webbrowser.Error: + rows.append(("browser", False, "no browser registered; --open will fail")) + rows.append(("colour", True if console.color_system else None, + console.color_system or "none - output is piped, or NO_COLOR is set")) + sd = state_dir() + writable = True + try: + sd.mkdir(parents=True, exist_ok=True) + probe = sd / ".probe" + probe.write_text("", encoding="utf-8") + probe.unlink() + except OSError: + writable = False + rows.append(("state dir", writable, str(sd))) + ack = load_ack() + rows.append(("authorisation", True if ack else None, + f"acknowledged {ack['at'][:10]}" if ack + else "not yet acknowledged - you will be asked on first run")) + rows.append(("python", True, platform.python_version())) + rows.append(("platform", True, f"{platform.system()} {platform.machine()}")) + + table = Table(box=box.SIMPLE_HEAD, border_style="df.rule", header_style="df.eyebrow", + pad_edge=False) + table.add_column("CHECK", style="df.text", no_wrap=True) + table.add_column("", justify="center", no_wrap=True) + table.add_column("DETAIL", style="df.dim") + for name, ok, detail in rows: + if ok is None: + mark = Text(g("info"), style="df.gold") + elif ok: + mark = Text(g("check"), style="df.ok") + else: + mark = Text(g("cross"), style="df.love") + table.add_row(name, mark, detail) + console.print(table) + console.print() + return 0 if all(ok is not False for _, ok, _ in rows) else 1 + + +# --------------------------------------------------------------------------- +# Interactive flow +# --------------------------------------------------------------------------- +def _ask(prompt): + """Run a questionary prompt, treating Ctrl-C as 'go back'.""" + try: + return prompt.ask() + except KeyboardInterrupt: + return None + + +def pick_objective() -> str | None: + choices = [Choice(title=f"{g(o.glyph)} {o.label}", value=o.key) for o in OBJECTIVES] + choices += [ + questionary.Separator(" " + "-" * 30), + Choice(title=f"{g('code')} Custom / freeform dork", value="__custom__"), + Choice(title=f"{g('book')} Operator reference", value="__ops__"), + Choice(title=f"{g('cross')} Quit", value="__quit__"), + ] + return _ask(questionary.select( + "What are you hunting for?", choices=choices, style=QUESTIONARY_STYLE, + qmark=g("search"), instruction=" ", + )) + + +def pick_engine(default: str = "google") -> Engine | None: + choices = [ + Choice(title=f"{e.label:<22}{e.notes.split('.')[0][:46]}", value=k) + for k, e in ENGINES.items() + ] + key = _ask(questionary.select( + "Which engine?", choices=choices, default=default, style=QUESTIONARY_STYLE, + qmark=g("globe"), instruction=" ", + )) + return ENGINES[key] if key else None + + +def ask_fields(needs: tuple[str, ...], preset: dict[str, str]) -> dict[str, str] | None: + vals = dict(PLACEHOLDERS) + for name in needs: + if preset.get(name): + vals[name] = preset[name] + continue + label, example, hint = FIELD_PROMPTS[name] + ans = _ask(questionary.text( + f"{label}:", style=QUESTIONARY_STYLE, qmark=g("arrow"), + instruction=f"({hint}, e.g. {example}) ", + validate=lambda t: True if t.strip() else "required", + )) + if ans is None: + return None + vals[name] = ans.strip() + return vals + + +def pick_recipe(console: Console, obj: Objective, vals: dict[str, str]) -> Recipe | str | None: + width = max(40, console.width - 14) + choices = [] + for i, r in enumerate(obj.recipes): + q = r.template.format(**vals).strip() + label = q if len(q) <= width else q[: width - 1] + "…" + choices.append(Choice(title=label, value=i)) + choices += [ + questionary.Separator(" " + "-" * 30), + Choice(title=f"{g('bolt')} Build all {len(obj.recipes)}", value="__all__"), + Choice(title=f"{g('arrow')} Back", value="__back__"), + ] + picked = _ask(questionary.select( + "Which dork?", choices=choices, style=QUESTIONARY_STYLE, qmark=g(obj.glyph), + instruction=" ", + )) + if picked is None or picked == "__back__": + return None + if picked == "__all__": + return "__all__" + return obj.recipes[picked] + + +def deliver(console: Console, engine: Engine, query: str, why: str, + session: SessionLog | None) -> str | None: + """Show a dork, explain it, then offer the action menu. Returns a signal.""" + translated, notes = translate(query, engine) + show_dork(console, engine, translated, why, notes) + explain(console, translated, engine) + + while True: + choices = [ + Choice(title=f"{g('copy')} Copy to clipboard", value="copy"), + Choice(title=f"{g('globe')} Open in browser", value="open"), + Choice(title=f"{g('save')} Save to session log", value="save"), + Choice(title=f"{g('book')} Explain an operator", value="explain"), + Choice(title=f"{g('bolt')} Refine this query", value="refine"), + questionary.Separator(" " + "-" * 30), + Choice(title=f"{g('arrow')} Another dork", value="back"), + Choice(title=f"{g('search')} New objective", value="objective"), + Choice(title=f"{g('cross')} Quit", value="quit"), + ] + action = _ask(questionary.select( + "Now what?", choices=choices, style=QUESTIONARY_STYLE, qmark=g("arrow"), + instruction=" ", + )) + if action in (None, "back"): + return "back" + if action in ("objective", "quit"): + return action + + if action == "copy": + ok, detail = copy_to_clipboard(translated) + icon, style = (g("check"), "df.ok") if ok else (g("cross"), "df.love") + console.print(Text(f" {icon} " + + (f"copied ({detail})" if ok else detail), style=style)) + elif action == "open": + url = search_url(engine, translated) + opened = webbrowser.open(url) + if opened: + console.print(Text(f" {g('check')} opened in your browser", style="df.ok")) + console.print(Text(f" {url}", style="df.faint")) + else: + console.print(Text(f" {g('cross')} could not open a browser. URL:", + style="df.love")) + console.print(Text(f" {url}", style="df.faint")) + elif action == "save": + if session is None: + console.print(Text(f" {g('cross')} no session log - start with " + f"--session FILE", style="df.love")) + else: + session.add(engine, translated, why, notes) + console.print(Text(f" {g('check')} appended to {session.path} " + f"(#{session.count})", style="df.ok")) + elif action == "explain": + name = _ask(questionary.text( + "Operator:", style=QUESTIONARY_STYLE, qmark=g("book"), + instruction="(e.g. site:, intitle:, AROUND) ", + )) + if name: + try: + operator_reference(console, name) + except SystemExit: + pass + elif action == "refine": + edited = _ask(questionary.text( + "Query:", default=translated, style=QUESTIONARY_STYLE, qmark=g("bolt"), + )) + if edited and edited.strip(): + translated = edited.strip() + _, notes = translate(translated, engine) + show_dork(console, engine, translated, why, notes) + explain(console, translated, engine) + + +def interactive(console: Console, args: argparse.Namespace, scope: str) -> int: + session = SessionLog(Path(args.session).expanduser(), scope) if args.session else None + if session: + console.print(Text(f" {g('save')} session log: {session.path}", style="df.faint")) + console.print() + preset = {"domain": args.domain or "", "org": args.org or "", + "keyword": args.keyword or "", "ext": args.ext or ""} + engine = ENGINES[args.engine] + + while True: + key = pick_objective() + if key in (None, "__quit__"): + break + if key == "__ops__": + operator_reference(console) + continue + if key == "__custom__": + chosen = pick_engine(engine.key) + if chosen is None: + continue + engine = chosen + raw = _ask(questionary.text( + "Your dork:", style=QUESTIONARY_STYLE, qmark=g("code"), + instruction="(Google dialect - it gets translated) ", + )) + if not raw or not raw.strip(): + continue + signal = deliver(console, engine, raw.strip(), "", session) + if signal == "quit": + break + continue + + obj = OBJECTIVE_BY_KEY[key] + console.print() + head = Text(" " + g(obj.glyph) + " ", style="df.foam") + head.append(obj.label, style="df.head") + console.print(head) + for line in textwrap.wrap(obj.blurb, width=min(86, console.width - 4)): + console.print(Text(" " + line, style="df.faint")) + if obj.caution: + console.print() + console.print(Text(f" {g('warn')} {obj.caution}", style="df.warn")) + console.print() + + vals = ask_fields(obj.needs, preset) + if vals is None: + continue + for f in obj.needs: + preset[f] = vals[f] + + chosen = pick_engine(engine.key) + if chosen is None: + continue + engine = chosen + + quit_now = False + while True: + recipe = pick_recipe(console, obj, vals) + if recipe is None: + break + if recipe == "__all__": + for r in obj.recipes: + q, notes = translate(r.template.format(**vals).strip(), engine) + show_dork(console, engine, q, r.why, notes) + if session: + session.add(engine, q, r.why, notes) + if session: + console.print() + console.print(Text(f" {g('check')} {len(obj.recipes)} dorks appended " + f"to {session.path}", style="df.ok")) + continue + signal = deliver(console, engine, recipe.template.format(**vals).strip(), + recipe.why, session) + if signal == "quit": + quit_now = True + break + if signal == "objective": + break + if quit_now: + break + + console.print() + if session and session.count: + console.print(Text(f" {g('save')} {session.count} dork(s) written to {session.path}", + style="df.ok")) + console.print(Text(f" {g('dot')} cheatsheet: {SHEET_URL}", style="df.faint")) + console.print() + return 0 + + +def oneshot(console: Console, args: argparse.Namespace, scope: str) -> int: + obj = OBJECTIVE_BY_KEY.get(args.objective) + if obj is None: + console.print(Text(f" unknown objective '{args.objective}'. Known: " + + ", ".join(OBJECTIVE_BY_KEY), style="df.bad")) + return 2 + vals = dict(PLACEHOLDERS) + for name in obj.needs: + supplied = getattr(args, name, None) + if not supplied: + console.print(Text(f" {g('warn')} --{name} not given; leaving {PLACEHOLDERS[name]} " + f"in the query", style="df.warn")) + else: + vals[name] = supplied + engine = ENGINES[args.engine] + session = SessionLog(Path(args.session).expanduser(), scope) if args.session else None + for r in obj.recipes: + q, notes = translate(r.template.format(**vals).strip(), engine) + show_dork(console, engine, q, r.why, notes) + if session: + session.add(engine, q, r.why, notes) + console.print() + if session: + console.print(Text(f" {g('check')} {session.count} dork(s) written to {session.path}", + style="df.ok")) + return 0 + + +# --------------------------------------------------------------------------- +# CLI +# --------------------------------------------------------------------------- +def build_parser() -> argparse.ArgumentParser: + p = argparse.ArgumentParser( + prog="dorkforge", + description="DAEMON//SEC search-operator workbench - build, understand " + "and translate search-engine dorks.", + epilog=f"cheatsheet: {SHEET_URL}", + formatter_class=argparse.RawDescriptionHelpFormatter, + ) + p.add_argument("--objective", metavar="KEY", + help="run non-interactively for one objective (see --list)") + p.add_argument("--domain", metavar="D", help="target domain, e.g. example.com") + p.add_argument("--org", metavar="O", help="organisation name") + p.add_argument("--keyword", metavar="K", help="hunt-specific keyword") + p.add_argument("--ext", metavar="E", help="file extension, without the dot") + p.add_argument("--engine", default="google", choices=sorted(ENGINES), + help="engine dialect to emit (default: google)") + p.add_argument("--session", metavar="FILE", + help="append every dork to this markdown log") + p.add_argument("--list", action="store_true", help="print the whole dork library and exit") + p.add_argument("--operators", action="store_true", + help="print the operator reference table and exit") + p.add_argument("--explain", metavar="OP", + help="explain one operator in full, e.g. --explain site:") + p.add_argument("--doctor", action="store_true", + help="check clipboard, browser, colour and state directory") + p.add_argument("--reset-ack", action="store_true", + help="forget the authorisation acknowledgement and ask again") + p.add_argument("--ascii", action="store_true", help="ASCII markers instead of Nerd Font glyphs") + p.add_argument("--no-color", action="store_true", help="disable colour output") + p.add_argument("--version", action="version", version=f"dorkforge {VERSION}") + return p + + +def main(argv: list[str] | None = None) -> int: + args = build_parser().parse_args(argv) + if args.ascii: + _G.clear() + _G.update(ASCII_GLYPHS) + console = Console(theme=THEME, no_color=args.no_color, highlight=False, soft_wrap=False) + + if args.doctor: + return doctor(console) + if args.reset_ack: + try: + ack_path().unlink() + console.print(Text(f" {g('check')} acknowledgement cleared", style="df.ok")) + except FileNotFoundError: + console.print(Text(f" {g('info')} nothing to clear", style="df.faint")) + except OSError as exc: + console.print(Text(f" {g('cross')} {exc}", style="df.love")) + return 1 + if not (args.objective or args.list or args.operators or args.explain): + return 0 + if args.explain: + banner(console) + operator_reference(console, args.explain) + return 0 + if args.operators: + banner(console) + operator_reference(console) + return 0 + if args.list: + banner(console) + library_dump(console) + return 0 + + banner(console) + + # Everything past this point wants a keyboard. --list/--operators/--explain + # and an already-acknowledged --objective run are the non-interactive paths. + needs_prompt = load_ack() is None or not args.objective + if needs_prompt and not sys.stdin.isatty(): + console.print(Text(f" {g('cross')} dorkforge is interactive and stdin is not a " + f"terminal.", style="df.bad")) + console.print(Text(" Run it in a terminal, or use the non-interactive paths:", + style="df.dim")) + console.print(Text(" dorkforge --list the whole dork library", + style="df.faint")) + console.print(Text(" dorkforge --operators the operator reference", + style="df.faint")) + console.print(Text(" dorkforge --objective files --domain example.com", + style="df.faint")) + console.print() + return 2 + + scope = gate(console) + if args.objective: + return oneshot(console, args, scope) + return interactive(console, args, scope) + + +if __name__ == "__main__": + try: + raise SystemExit(main()) + except KeyboardInterrupt: + print() + raise SystemExit(130) diff --git a/public/downloads/enumeration/dorkforge.py.sha256 b/public/downloads/enumeration/dorkforge.py.sha256 @@ -0,0 +1 @@ +3173df88ab3102ebc4999985fcb881298f3f51374170592e5ac95d0674210193 dorkforge.py diff --git a/src/content/sheets/enumeration/google-dorking.md b/src/content/sheets/enumeration/google-dorking.md @@ -0,0 +1,583 @@ +--- +title: "Google Dorking" +description: "Search-engine operators as a recon primitive: the full operator reference, the gotchas that silently break dorks, recipes by objective, cross-engine translation, OPSEC, and defending your own estate against them." +category: enumeration +tags: [enumeration, osint, recon, google-dorking, dorks, passive, search-operators, ghdb] +tools: [Google, Bing, DuckDuckGo, Yandex, GitHub, Shodan, dorkforge] +difficulty: beginner +updated: "2026-09-26" +--- + +# Google Dorking + +Dorking is using a search engine's own query language to find things its index already holds but nobody meant to publish: config files, backups, admin panels, stack traces, cloud buckets, private keys. No packet reaches the target — you are querying Google's copy of the web, which is why it opens [Stage 00 — Passive External Recon](/sheets/pentest-workflow/passive-external-recon) and why it survives every firewall, WAF and rate limit the target owns. + +The technique is old and the operators are few. What separates a useful dork from noise is knowing which operators still work, how they combine, and where each one silently fails. + +> [!danger] Indexed is not authorised +> Every dork on this page returns *links*. Following one to retrieve a config file, enumerate a bucket, or log into a panel is **access**, and access needs written permission — a public URL is not consent, and "it was in Google" has never been a defence. Personal data you collect is regulated (GDPR and equivalents) whether or not the engagement covers it. Use this against your own estate, a scope you hold in writing, or a lab you are entitled to. + +--- + +## 1. Why it works — you are searching the index, not the site + +A crawler finds a URL, fetches it, and the indexer stores the text, title, URL and file type. Everything dorking does is filter that stored record. Three consequences matter: + +| Fact | Consequence for recon | +|---|---| +| `robots.txt` is a **crawl** directive, not an access control | A `Disallow:` line stops polite crawlers fetching a path — it does not stop the path being served, and it publishes a list of the paths the owner considers sensitive. Fetch `$DOMAIN/robots.txt` first; it is a free sitemap of the interesting bits. | +| Only `noindex` removes a page from results | `X-Robots-Tag: noindex` or a `<meta name="robots" content="noindex">` tag is the actual removal mechanism, and it only works if the crawler is *allowed in* to see it. A `Disallow`ed page that is linked from elsewhere can still appear. | +| The index outlives the file | Deleting a file does not delete the record. Titles and snippets persist for weeks; archives persist indefinitely. A dork that returns a 404 still tells you the file existed and what it was called. | +| Anything linked once is crawlable | A "secret" URL pasted into a public ticket, a README, a Slack export or a CDN log gets crawled. This is how private buckets and staging hosts end up indexed. | + +The practical reading: **the index is a historical record of what the target served, not a snapshot of what it serves now.** Treat every hit as a lead with a timestamp, not a live finding. + +--- + +## 2. Anatomy of a dork + +```text +site:example.com ext:sql intext:"CREATE TABLE" -site:docs.example.com +└──── scope ───┘ └─ type ┘ └───── content ────┘ └────── exclusion ─────┘ +``` + +Four moving parts, and only the first is mandatory in practice: + +- **Scope** — which hosts count (`site:`). +- **Type** — what kind of document (`ext:` / `filetype:`). +- **Content** — a literal that only the interesting version of the file contains (`intext:`, `intitle:`, `inurl:`). +- **Exclusion** — what to subtract once the first pass is noisy (`-site:`, `-inurl:`). + +Terms and operators are **AND**-ed by default. There is no `AND` keyword — writing one searches for the word "and". + +> [!tip] Build them in that order +> Scope first, then type, then narrow with content, then subtract noise. Starting from a long stacked dork you copied off a list gives you zero results and no idea which clause killed it. Start broad, add one operator at a time, and watch the result count move. + +--- + +## 3. Operator reference + +### Scope + +| Operator | Does | Example | +|---|---|---| +| `site:` | Restrict to a host. Matches the host as a **suffix**, so it covers subdomains. | `site:example.com` | +| `site:*.` | Force the subdomain tree. Pair with a negative to drop the ones you know. | `site:*.example.com -site:www.example.com` | +| `-site:` | Subtract a host. The main tool for forcing new subdomains to the surface. | `-site:blog.example.com` | +| `related:` | Sites Google considers similar. **Officially unsupported** — usually returns nothing. | `related:example.com` | + +### Content position + +| Operator | Matches in | Notes | +|---|---|---| +| `intitle:` | The `<title>` element | Highest-signal field for fingerprinting an app by its shipped default title. | +| `inurl:` | The whole URL, **including the query string** | This is why it finds parameter names: `inurl:redirect=`. Substring match. | +| `intext:` | The page body | Forces a term into the body instead of letting it match the title or an inbound link. | +| `inanchor:` | Inbound link text | Thin and unreliable now that link signals are de-emphasised. | +| `allintitle:` `allinurl:` `allintext:` | Same fields, all following terms | **Do not use.** See the gotchas below. | + +### File type + +| Operator | Does | Notes | +|---|---|---| +| `filetype:` | One file extension | Matches the extension Google recorded, not the `Content-Type` header. | +| `ext:` | Identical to `filetype:` | Shorter, so stacked dorks stay readable. Both work on Google, Bing and Brave. | + +One extension per operator — stack them with `OR` to cover a family: + +```text +site:$DOMAIN ext:sql OR ext:bak OR ext:old OR ext:backup +``` + +Google indexes a fixed set of document types (`pdf`, `doc`/`docx`, `xls`/`xlsx`, `ppt`/`pptx`, `rtf`, `ps`, `kml`, `txt` and a handful more). A `ext:env` dork works **only because such files were served as plain text and indexed as text** — which is precisely what makes a hit worth reporting. + +### Logic and grouping + +| Syntax | Does | Trap | +|---|---|---| +| `"phrase"` | Exact match, in order, no stemming or synonyms | Any literal artefact — a banner, an error string, a key header — belongs in quotes. | +| `OR` or `\|` | Disjunction | **Must be uppercase.** Lowercase `or` is a stop word and is silently ignored. | +| `-term` | Exclude | No space after the hyphen. Works on operators too: `-site:`, `-inurl:`. | +| `( )` | Grouping | Required whenever you mix `OR` with AND-ed terms, or precedence bites you. | +| `*` | Whole-token wildcard | Only meaningful **inside quotes**. It is not a substring wildcard — `adm*n` does not work. | +| `AROUND(n)` | The two terms within *n* words | Uppercase, no space before the bracket. Tighter than AND, looser than a phrase. | + +```text +site:$DOMAIN (ext:sql OR ext:bak) intext:password # grouped correctly +site:$DOMAIN ext:sql OR ext:bak intext:password # ambiguous — do not +password AROUND(3) database site:$DOMAIN # proximity +"confidential * report" site:$DOMAIN # token wildcard +``` + +### Time and number + +| Syntax | Does | Caveat | +|---|---|---| +| `before:YYYY-MM-DD` | Documents Google dates before that day | Filters on Google's *estimate* of the document date, which is frequently wrong for pages without date markup. | +| `after:YYYY-MM-DD` | Documents Google dates after that day | Same caveat. Good for narrowing; never good enough to prove publication date. | +| `1000..2000` | Any number in the range | Two dots. Legal inside other operators: `inurl:id=1..100`. | + +### Retired — and what replaced them + +> [!warning] These appear in every dork list on the internet and none of them work +> | Operator | Status | Use instead | +> |---|---|---| +> | `cache:` | Removed by Google in **September 2024**, along with the cached-page links | [web.archive.org](https://web.archive.org), [archive.today](https://archive.today) | +> | `link:` | Deprecated **2017**; now a silent no-op that degrades to a keyword search | Search Console (for sites you own) or a commercial backlink tool | +> | `info:` | Retired; folded into ordinary search | Just search the bare URL | +> | `+term` | Removed **2011** | Quote the term instead: `"term"` | +> | `daterange:` | Julian-date legacy syntax, unreliable | `before:` / `after:` | +> +> A retired operator does not error. Google drops it and searches the remaining words, so your dork returns plausible-looking garbage. This is the single most common reason a copied dork "stops working". + +--- + +## 4. The gotchas that actually break dorks + +> [!example] Read this section twice — it is most of the skill +> | Trap | What happens | Fix | +> |---|---|---| +> | Lowercase `or` | Treated as a stop word, dropped. Your disjunction became an AND. | Always uppercase `OR`. | +> | `allintitle:` / `allinurl:` / `allintext:` | They consume **the rest of the query** — every operator after them is ignored, including `site:`. Your scoped dork quietly went global. | Stack the singular form instead: `intitle:a intitle:b`. | +> | Space after the colon | `site: example.com` is parsed as a bare `site:` plus the keyword `example.com`. | Never put a space after an operator's colon. | +> | Unquoted multi-word value | `intitle:index of` means `intitle:index` AND the word `of`. | Quote it: `intitle:"index of"`. | +> | `*` outside quotes | Ignored entirely. | Only use it inside a quoted phrase. | +> | `site:` is a suffix match | `site:example.com` also returns `notexample.com`-style matches on some engines and always includes every subdomain. | Add `-site:` exclusions, or use Yandex's stricter `host:`. | +> | Very long stacked dorks | Google historically truncated long queries, and still degrades them — the tail of a 15-operator dork may never be evaluated. | Split into two or three narrower dorks. | +> | Personalised results | Signed in, Google reorders and filters results for *you*. Two operators can look broken because your profile suppressed the hits. | Search signed out, in a clean profile. | +> | Result count is a lie | The "About N results" figure is an estimate and swings wildly. | Page to the end. The real count is where results stop. | +> | Verbatim mode | Google rewrites and expands queries by default. | Append `&tbs=li:1` to the results URL, or use Tools → All results → Verbatim. | + +> [!tip] Two switches worth memorising +> `&num=100` on the results URL returns 100 per page instead of 10. `&filter=0` disables the "omitted similar results" collapse, which routinely hides the interesting duplicates on a large site. + +--- + +## 5. Recipes by objective + +Every dork below is written in the Google dialect with `$DOMAIN`, `$ORG` and `$KEYWORD` as placeholders. Section 6 covers translating them. + +### Exposed files and configs + +The highest-yield category in the whole practice. + +| Dork | Finds | +|---|---| +| `site:$DOMAIN ext:env OR ext:cfg OR ext:conf OR ext:ini` | Application config. A served `.env` is plaintext credentials: database URI, mail password, cloud keys, app secret. | +| `site:$DOMAIN ext:bak OR ext:old OR ext:backup OR ext:swp` | Editor and deploy leftovers. `index.php.bak` is served as text instead of executed — that is the source code. | +| `site:$DOMAIN ext:log OR ext:txt inurl:log` | Logs: internal paths, usernames, session tokens, stack traces. | +| `site:$DOMAIN ext:yml OR ext:yaml OR ext:toml inurl:config` | Compose files, CI definitions, `appsettings`, secrets baked into a values file. | +| `site:$DOMAIN ext:pem OR ext:key OR ext:ppk OR ext:p12 OR ext:pfx` | Private keys and cert bundles. Any hit is a critical finding on its own. | +| `site:$DOMAIN inurl:.git OR inurl:.svn OR inurl:.hg` | Exposed VCS metadata — a readable `.git/` means the whole repo and its history. | +| `site:$DOMAIN inurl:wp-config OR inurl:configuration.php OR inurl:settings.py` | Framework config by canonical filename. | + +### Login and admin panels + +| Dork | Finds | +|---|---| +| `site:$DOMAIN inurl:admin OR inurl:administrator OR inurl:adminpanel` | The obvious paths. Still hits constantly. | +| `site:$DOMAIN intitle:"login" OR intitle:"sign in"` | Title-based, so it catches panels on paths you would never guess. | +| `site:$DOMAIN inurl:phpmyadmin OR inurl:adminer OR inurl:pma` | Exposed database consoles, frequently on vendor defaults. | +| `site:$DOMAIN intitle:"Dashboard" (inurl:jenkins OR inurl:grafana OR inurl:kibana)` | CI and observability — build secrets, internal hostnames, production telemetry. | +| `site:$DOMAIN inurl:/manager/html OR inurl:/jmx-console OR inurl:/axis2` | Java app-server managers. Tomcat manager with default creds is a WAR deploy from RCE. | +| `site:$DOMAIN inurl:owa OR inurl:/rdweb OR inurl:citrix OR inurl:vpn` | Remote-access front doors — the surface spraying actually targets. | + +### Open directory listings + +| Dork | Finds | +|---|---| +| `site:$DOMAIN intitle:"index of"` | The canonical dork. Apache, nginx and IIS all emit a title starting this way. | +| `site:$DOMAIN intitle:"index of" (backup OR bak OR old OR archive)` | Listings that name themselves as a dumping ground. | +| `site:$DOMAIN intitle:"index of" "parent directory" (.sql OR .zip OR .tar.gz)` | Listings with a dump or archive sitting in them. | +| `site:$DOMAIN "Directory Listing For" -inurl:html` | Tomcat and others whose listing wording differs from Apache's. | + +### Cloud storage + +| Dork | Finds | +|---|---| +| `site:s3.amazonaws.com "$ORG"` | S3 buckets whose listing mentions the organisation. | +| `site:$DOMAIN inurl:s3.amazonaws.com` | Pages on the target that link into S3 — the fastest way to learn the bucket naming convention. | +| `inurl:blob.core.windows.net "$ORG"` | Azure Blob containers. | +| `inurl:storage.googleapis.com "$ORG"` | Google Cloud Storage objects. | +| `site:$DOMAIN inurl:r2.dev OR inurl:digitaloceanspaces.com` | Smaller providers — less scrutiny, misconfigured more often. | +| `"$ORG" (site:trello.com OR site:notion.site)` | Public boards and docs. Trello leaks credentials and architecture diagrams at a remarkable rate. | + +### Credentials, keys and tokens + +| Dork | Finds | +|---|---| +| `site:$DOMAIN intext:"BEGIN RSA PRIVATE KEY"` | Key headers pasted into a wiki, a ticket, a README. | +| `site:$DOMAIN intext:"AKIA"` | AWS access key IDs all begin `AKIA` — a distinctive literal that rarely false-positives. | +| `"$ORG" ("xoxb-" OR "ghp_" OR "sk_live_")` | Slack, GitHub and Stripe token prefixes. Structurally unique, so a hit is almost never noise. | +| `site:$DOMAIN (ext:env OR ext:yml) intext:PASSWORD` | Config files with a password assignment. | +| `site:pastebin.com OR site:justpaste.it "$DOMAIN"` | Paste sites — where dumps, configs and credential lists get parked. | + +> [!danger] A live key is a finding, not a login +> Do not authenticate with anything you find. Record the URL, report it immediately so it can be rotated, and treat your own notes as sensitive material for the rest of the engagement. See [Credential Hunting](/sheets/enumeration/credential-hunting), [Gitleaks](/sheets/enumeration/2-4-cheatsheet-gitleaks) and [TruffleHog](/sheets/enumeration/2-5-cheatsheet-trufflehog) for the systematic version of this. + +### Documents and metadata + +| Dork | Finds | +|---|---| +| `site:$DOMAIN ext:pdf OR ext:docx OR ext:xlsx OR ext:pptx` | The corpus. Pull it, then run `exiftool` over the lot. | +| `site:$DOMAIN ext:pdf (confidential OR internal OR "not for distribution")` | Documents labelled restricted that are nonetheless public. | +| `site:$DOMAIN ext:xlsx (employee OR salary OR roster)` | Spreadsheets — the format people paste staff lists into. | +| `site:$DOMAIN ext:vsd OR ext:vsdx OR ext:drawio` | Network and architecture diagrams. | + +```bash +# the half of this that isn't the dork: metadata gives you usernames and internal paths +exiftool -Author -Creator -Producer -Company *.pdf | sort -u +exiftool -a -G1 -s report.docx | grep -Ei 'author|company|template|lastmodifiedby' +``` + +### Error messages and stack traces + +| Dork | Finds | +|---|---| +| `site:$DOMAIN intext:"You have an error in your SQL syntax"` | The literal MySQL parser message — confirms injectable input. | +| `site:$DOMAIN intext:"Traceback (most recent call last)"` | Python tracebacks. Django debug pages additionally dump settings and environment. | +| `site:$DOMAIN intitle:"Whoops, looks like something went wrong"` | Laravel's debug page, which shows environment variables verbatim. | +| `site:$DOMAIN intext:"Fatal error" OR intext:"Uncaught exception"` | PHP and Java fatals — usually carry the absolute filesystem path. | +| `site:$DOMAIN intext:"Server Error in" intext:"Stack Trace"` | ASP.NET yellow-screen-of-death with tracing left on. | +| `site:$DOMAIN intitle:"phpinfo()"` | Full build config, loaded modules, absolute paths, sometimes environment secrets. | + +### Databases, backups and dumps + +| Dork | Finds | +|---|---| +| `site:$DOMAIN ext:sql intext:"INSERT INTO"` | SQL dumps confirmed by their own statement syntax, not just an extension. | +| `site:$DOMAIN ext:sql intext:"CREATE TABLE" intext:password` | Dumps that include a credentials table. | +| `site:$DOMAIN (ext:zip OR ext:tar OR ext:gz) (backup OR dump OR export)` | Archived backups in the webroot. | +| `site:$DOMAIN inurl:backup OR inurl:dump OR inurl:export` | Path-based — catches the directory even when the file was not indexed. | + +### Injectable parameters + +A target list for later, authorised testing. Nothing here is a finding by itself. + +| Dork | Vulnerability class | +|---|---| +| `site:$DOMAIN inurl:id= OR inurl:pid= OR inurl:cat=` | SQL injection, IDOR | +| `site:$DOMAIN inurl:file= OR inurl:page= OR inurl:path=` | LFI, path traversal — see [LFI](/sheets/enumeration/lfi) | +| `site:$DOMAIN inurl:url= OR inurl:redirect= OR inurl:next=` | Open redirect, the front half of many SSRF chains | +| `site:$DOMAIN inurl:cmd= OR inurl:exec= OR inurl:query=` | Command injection, reflected XSS | +| `site:$DOMAIN inurl:debug=true OR inurl:admin=1` | Flags that flip the app into a verbose or privileged mode | + +### People, emails and usernames + +| Dork | Finds | +|---|---| +| `site:linkedin.com/in "$ORG"` | Staff profiles — names plus role, which is what derives the username scheme. | +| `site:$DOMAIN intext:"@$DOMAIN"` | The org's own pages leaking its address format (`first.last@`, `flast@`). | +| `"$ORG" (site:github.com OR site:gitlab.com) intext:"@$DOMAIN"` | Developers using a work address in commits or profiles. | +| `"$ORG" (site:stackoverflow.com OR site:serverfault.com)` | Engineers asking about their own stack, pasting real config with real hostnames. | +| `"$ORG" ("we use" OR "powered by") -site:$DOMAIN` | Tech stack from job ads, case studies, conference talks. | + +> [!warning] Personal data is regulated independently of your scope +> Collect the minimum that supports the finding, store it with the engagement material, and delete it when the engagement closes. + +### Subdomains and forgotten hosts + +| Dork | Finds | +|---|---| +| `site:*.$DOMAIN -site:www.$DOMAIN` | The core subdomain dork. | +| `site:*.$DOMAIN (dev OR staging OR test OR uat OR qa)` | Non-production — weaker credentials, debug on, real data. | +| `site:*.$DOMAIN (vpn OR remote OR portal OR intranet)` | Access infrastructure and internal-facing hosts. | +| `"$DOMAIN" -site:$DOMAIN` | Third-party references: partners, status pages, monitoring, archived copies. | + +Dorking is the *weakest* of the subdomain sources — use it to catch what certificate transparency and DNS missed, not as the primary pass. See [Amass](/sheets/enumeration/amass) and [Stage 00](/sheets/pentest-workflow/passive-external-recon) for the real enumeration. + +### Archived and removed pages + +| Dork | Finds | +|---|---| +| `site:web.archive.org "$DOMAIN"` | Wayback captures Google indexed. | +| `site:archive.today "$DOMAIN"` | The other archive, which captures pages Wayback refuses. | +| `site:$DOMAIN before:2022-01-01 "$KEYWORD"` | Older material still on the live site. | + +```bash +# the archive's own API beats dorking it — every captured URL, deduplicated +curl -s "http://web.archive.org/cdx/search/cdx?url=$DOMAIN*&output=text&fl=original&collapse=urlkey" \ + | sort -u > wayback-urls.txt +# pull the ones that look like config or secrets out of the pile +grep -Ei '\.(env|sql|bak|old|log|yml|ini|conf)($|\?)' wayback-urls.txt +``` + +--- + +## 6. Cross-engine translation + +Google is the reference dialect, not the only one. Yandex and Bing routinely index hosts Google never crawled, so a subdomain dork is worth running through at least two engines. + +| Intent | Google | Bing | DuckDuckGo | Yandex | Brave | +|---|---|---|---|---|---| +| Scope to host | `site:` | `site:` | `site:` | `site:` / `host:` / `rhost:` | `site:` | +| File type | `filetype:` `ext:` | `filetype:` `ext:` | `filetype:` | `mime:` | `filetype:` | +| In title | `intitle:` | `intitle:` | `intitle:` | `title:` | `intitle:` | +| In URL | `inurl:` | `inurl:` | `inurl:` | `url:` | `inurl:` | +| In body | `intext:` | `inbody:` | — | *(default)* | — | +| Exact phrase | `"..."` | `"..."` | `"..."` | `"..."` | `"..."` | +| Exclude | `-term` | `-term` | `-term` | `-term` | `-term` | +| Either/or | `OR` | `OR` | `OR` | `OR` | `OR` | +| Proximity | `AROUND(n)` | — | — | `/+n` | — | +| Date bound | `before:` `after:` | *(UI filter)* | *(UI filter)* | `date:` | *(UI filter)* | + +**Engine-native operators with no Google equivalent — worth the detour on their own:** + +| Engine | Operator | Why you would bother | +|---|---|---| +| Bing | `ip:203.0.113.10` | Every site Bing indexed on that address — a free reverse-IP lookup, and often the rest of a shared estate behind one origin. | +| Bing | `contains:sql` | Pages that **link to** a `.sql` file rather than the file itself — catches the index page of a backup directory that was never indexed. | +| Bing | `url:example.com/path` | Tests whether one exact URL is in the index. | +| Yandex | `rhost:com.example.*` | Reversed-domain match across the whole tree — stricter and more predictable than Google's suffix matching. | +| Yandex | `mime:pdf` | Yandex's `filetype:`, keyed on recorded MIME type. | +| DuckDuckGo | `!bangs` (`!gh`, `!so`) | Not operators — redirects to another site's own search. Handy, but the query leaves DDG. | + +> [!note] Brave and DuckDuckGo are the low-friction options +> Neither requires an account and neither personalises results, which makes them the right place to start from a clean browser. Their operator support is genuinely narrower than Google's and less documented — when a dork returns nothing there, re-run it on Google before concluding the target is clean. + +--- + +## 7. Other dialects worth knowing + +### GitHub code search + +A different language entirely, and the single richest source of leaked secrets. Requires being signed in — use a throwaway account. + +| Operator | Example | Notes | +|---|---|---| +| `org:` | `org:$ORG path:.env` | Scope to an organisation's public repos. | +| `repo:` | `repo:owner/name "password"` | One repository. | +| `path:` | `path:.github/workflows` | Replaced the old `filename:`. Globs: `path:**/config/*.yml`. | +| `language:` | `"$DOMAIN" language:yaml` | Narrows a noisy content search. | +| `symbol:` | `symbol:connectDatabase` | Definitions rather than mentions. | + +```text +org:$ORG path:.env +org:$ORG "BEGIN RSA PRIVATE KEY" +org:$ORG path:.github/workflows +"$DOMAIN" ("192.168." OR "10.0." OR ".local") +"$DOMAIN" "password" language:yaml +``` + +The last two are the interesting ones: they find *other people's* repos that leak the target's internal addressing and credentials — contractors, ex-employees, integration partners. + +### Shodan and Censys + +These index service banners, not web pages, so the web operators do not apply at all. Covered properly in [Shodan](/sheets/enumeration/shodan); the translation for a dorking reflex: + +| You want | Shodan | +|---|---| +| Hosts presenting the org's certificate | `ssl.cert.subject.cn:$DOMAIN` | +| Pages referencing the domain | `http.html:"$DOMAIN"` | +| A specific page title | `http.title:"index of"` | +| The org's netblocks | `org:"$ORG"` or `net:203.0.113.0/24` | + +The origin-IP hunt behind a CDN is `ssl.cert.subject.cn:` — the origin still presents the real certificate. + +--- + +## 8. The Google Hacking Database + +The [GHDB](https://www.exploit-db.com/google-hacking-database) is Exploit-DB's curated archive of dorks, started by Johnny Long, now thousands of entries across categories: files containing passwords, sensitive directories, vulnerable servers, error messages, footholds. + +How to use it without wasting an afternoon: + +- **Filter by date.** Most of the archive is a decade old and targets software nobody runs. Sort newest first. +- **Read it as a pattern library, not a copy-paste list.** The value of `intitle:"index of" "service.pwd"` is not that exact string — it is that vendor-default filenames are the way in. Adapt the shape to the stack you actually found. +- **Scope everything.** Almost every GHDB entry is unscoped and returns strangers' infrastructure. Add `site:$DOMAIN` before you run it, every time. + +--- + +## 9. OPSEC + +> [!danger] You are the one being logged +> Dorking sends nothing to the target — but it sends everything to the search engine, tied to your account, your IP and your browser fingerprint. The queries themselves describe your intent with unusual clarity. + +| Rule | Why | +|---|---| +| Never dork signed in to a real account | The query history is retained, attributable, and personalises your results into uselessness. Use a clean profile or a private window. | +| Use a throwaway for GitHub code search | It is the one dialect that mandates an account. Do not attach your identity to a secrets hunt. | +| Expect CAPTCHAs, and slow down | A burst of operator-heavy queries from one IP trips Google's automation detection. Once you are CAPTCHA'd, that IP is degraded for hours. | +| Do not click through to live secrets | Fetching the `.env` is a request to the target — logged, attributable, and outside "passive" recon. The snippet in the result usually contains enough to report. | +| Treat your notes as sensitive | A session log full of working dorks and hit URLs is a map of the target's exposure. Store it with the engagement, encrypt it, delete it on close. | +| Report live keys immediately | Rotation is time-sensitive. Do not sit on a working credential until the report. | + +--- + +## 10. Automation, rate limits and the ToS question + +Scripted scraping of Google results violates its Terms of Service and fails in practice: the automation detection is good, the HTML changes, and the CAPTCHA is unsolvable at scale without paying someone to solve it. Every "google dork scanner" on GitHub has the same lifecycle — works for a week, then returns empty pages forever. + +The options that actually work: + +| Approach | Reality | +|---|---| +| **Do it by hand** | Slow, unblockable, and you read each result properly. For a single engagement this is genuinely the right answer. | +| **Programmable Search Engine JSON API** | Google's sanctioned path. Free tier is ~100 queries/day, then paid. Supports most operators. Results are scoped to a custom engine you configure, which can be set to search the whole web. | +| **A commercial SERP API** | Pays someone else to absorb the blocking. Fine for volume; check that sending client data to a third party is acceptable under your engagement's confidentiality terms. | +| **Shodan / Censys APIs** | Purpose-built, properly documented, generous limits. If the thing you want is a service rather than a document, stop dorking and query these. | +| **Bing Search APIs** | Microsoft has been winding these down — verify current availability before building anything on them. | + +A tool that *builds* dorks and hands them to your browser has none of these problems, because you are still the one searching. That is what [dorkforge](#13-dorkforge--the-companion-script) does. + +--- + +## 11. Blue team — defending your own estate + +Dorking is trivially turned around: run it against yourself, on a schedule. + +| Control | What it actually does | +|---|---| +| `X-Robots-Tag: noindex` header, or `<meta name="robots" content="noindex">` | **The** removal mechanism. Note the trap: a path blocked in `robots.txt` cannot be crawled, so the crawler never sees the `noindex` — the page can stay indexed forever. Allow the crawl, serve the `noindex`, then block it once it has dropped out. | +| Search Console removal tool | Fast takedown for a site you own (~6 months), which buys time while you fix the underlying exposure. It is a suppression, not a fix. | +| `robots.txt` | A crawl hint, and a published list of your sensitive paths. Never treat it as access control; assume attackers read it first. | +| Disable autoindex | `Options -Indexes` (Apache), `autoindex off` (nginx). Kills the entire `intitle:"index of"` class. | +| Block dotfiles and backup extensions at the edge | Return 404 for `\.(env\|git\|bak\|old\|sql\|swp)$` before the request reaches the app. | +| Turn off debug in production | `APP_DEBUG=false`, `customErrors="On"`, `DEBUG = False`. Removes the entire error-message class in one change. | +| Bucket ACLs and public-access blocks | Cloud storage exposure is a configuration problem; no search-engine control fixes it. | +| Continuous secret scanning | Gitleaks/TruffleHog in CI plus GitHub push protection catches the commit before it is ever indexed. | +| Dork yourself on a schedule | The blue-team use of this whole page. Automate the handful of dorks that matter for your estate and alert on new hits. | + +```bash +# a minimal self-audit, run monthly against your own domain +for d in 'ext:env' 'ext:sql' 'ext:bak' 'intitle:"index of"' 'intext:"BEGIN RSA PRIVATE KEY"'; do + printf '%s\n' "site:$DOMAIN $d" +done +# feed them to the Programmable Search JSON API and diff the hit list against last month's +``` + +--- + +## 12. Practising safely + +- **Your own domains.** The only fully unambiguous target, and the one where findings are actionable. +- **`google-gruyere.appspot.com`**, **`testphp.vulnweb.com`**, **`demo.testfire.net`** — deliberately vulnerable hosts published for training. Verify the terms on each before touching it. +- **The GHDB, read but not run.** Study the shapes; run them scoped to yourself. +- **CTF and lab platforms** with an OSINT category, where the scope is explicit. + +> [!warning] Unscoped dorks return strangers' infrastructure +> A dork like `intitle:"webcamXP"` returns hardware belonging to people who never consented to anything. Looking at the result list is one thing; connecting to a device is unauthorised access in essentially every jurisdiction. Scope with `site:` or do not run it. + +--- + +## 13. dorkforge — the companion script + +**[dorkforge.py](/downloads/enumeration/dorkforge.py)** ([SHA-256](/downloads/enumeration/dorkforge.py.sha256)) is a single-file interactive workbench for everything on this page: it asks what you are hunting for, builds the dork, **explains every operator it used**, translates it for the engine you picked, and hands it to your clipboard, your browser, or an engagement log. + +It ships the full library from section 5 — **14 objectives, 81 recipes, 32 operators** — and it sends no traffic to any target. The searching still happens in your browser, under your account, from your IP, which is why it opens with the authorisation gate and records a scope reference in every session log. + +### Run it + +It uses [PEP 723](https://peps.python.org/pep-0723/) inline metadata, so `uv` resolves `rich` and `questionary` into a throwaway environment on first run. Nothing to install, nothing to clean up. + +```bash +# download + verify +curl -sO https://cheatsheet.daemon-sec.xyz/downloads/enumeration/dorkforge.py +curl -sO https://cheatsheet.daemon-sec.xyz/downloads/enumeration/dorkforge.py.sha256 +shasum -a 256 -c dorkforge.py.sha256 # expect: dorkforge.py: OK + +# run it +uv run dorkforge.py +``` + +Expected SHA-256: + +```text +3173df88ab3102ebc4999985fcb881298f3f51374170592e5ac95d0674210193 +``` + +Or make it executable and let the shebang do the work — `#!/usr/bin/env -S uv run --script`: + +```bash +chmod +x dorkforge.py +./dorkforge.py +# keep it on PATH like the rest of your toolkit +install -m 755 dorkforge.py ~/.local/bin/dorkforge +``` + +> [!note] Don't have uv? +> `curl -LsSf https://astral.sh/uv/install.sh | sh` (or `brew install uv`). Failing that, `pip install rich questionary` and run it with plain `python3` — the script degrades to that path and tells you so. + +### The flow + +```text + ▪ DÆMON//SEC + dorkforge v1.0.0 search-operator workbench +──────────────────────────────────────────────────────────── + +? What are you hunting for Exposed files & configs +? Target domain: example.com +? Which engine? Google +? Which dork? site:example.com ext:env OR ext:cfg OR ext:conf OR ext:ini + +┌─ Google ─────────────────────────────────────────────────┐ +│ │ +│ site:example.com ext:env OR ext:cfg OR ext:conf OR ext:ini│ +│ │ +│ Application config. A served .env is credentials in │ +│ plaintext - database URI, mail password, cloud keys. │ +│ │ +└────────────────────────────────────────────────────────────┘ + + OPERATOR WHAT IT DOES GOOGLE + ────────────────────────────────────────────────────────────── + site: Only return pages whose host matches. ✓ + ext: Synonym for filetype: on Google and Bing. ✓ + OR Match either side. Must be uppercase. ✓ + +? Now what? Copy to clipboard / Open in browser / Save to log +``` + +Pick a non-Google engine and it rewrites the dork and tells you what changed: + +```text +$ uv run dorkforge.py --objective listing --domain example.com --engine yandex + + site:example.com title:"index of" + dialect notes + ▪ intitle: rewritten to title: for Yandex +``` + +### Usage + +| Command | Does | +|---|---| +| `uv run dorkforge.py` | Interactive. The main path. | +| `dorkforge --session engagement.md` | Append every dork you build to a markdown log, with the URL, the reasoning and the translation notes. | +| `dorkforge --objective files --domain example.com` | Non-interactive: print every recipe for one objective. | +| `dorkforge --objective code --org acme --engine github` | Same, in GitHub's dialect. | +| `dorkforge --list` | The whole library — 14 objectives, 81 recipes — to a pager or a file. | +| `dorkforge --operators` | The operator reference table, with live / unreliable / retired status. | +| `dorkforge --explain site:` | The long form on any one operator, plus which engines support it. | +| `dorkforge --doctor` | Check clipboard helper, browser, colour support and state directory. | +| `dorkforge --ascii` | ASCII markers instead of Nerd Font glyphs. | +| `dorkforge --reset-ack` | Forget the authorisation acknowledgement and ask again. | + +**Flags:** `--domain` `--org` `--keyword` `--ext` pre-fill the scope prompts. `--engine` picks the dialect (`google` `bing` `ddg` `yandex` `brave` `github` `shodan`). `--no-color` for piping. + +### The session log + +`--session` writes engagement-ready markdown as you work — the dork, why it was built, the live URL, and anything that did not survive translation: + +~~~markdown +# dorkforge session + +- **Started:** 2026-09-26 14:21 BST +- **Authorised scope:** ACME-2026-114 +- **Tool:** dorkforge 1.0.0 + +## 1. Google + +Application config. A served .env is credentials in plaintext. + +```text +site:example.com ext:env OR ext:cfg OR ext:conf OR ext:ini +``` + +<https://www.google.com/search?q=site%3Aexample.com+ext%3Aenv...> +~~~ + +Paste it straight into the recon section of a report — see [Documentation and Reporting](/sheets/pentest-workflow/documentation-and-reporting). + +> [!tip] The authorisation gate is once per machine +> The acknowledgement is stored in `$XDG_STATE_HOME/dorkforge/ack.json` (`~/.local/state/dorkforge/` by default) along with the scope reference you gave, which is then stamped into every session log. `--reset-ack` clears it. diff --git a/src/content/sheets/pentest-workflow/passive-external-recon.md b/src/content/sheets/pentest-workflow/passive-external-recon.md @@ -116,6 +116,8 @@ dig -x $IP # reverse DNS — DCs and mail servers oft ## 4. Search-engine dorks +The table below is the field-expedient subset. The operator reference, the traps that silently break a dork, per-objective recipes and the cross-engine translation table live on [Google Dorking](/sheets/enumeration/google-dorking) — which also ships [dorkforge](/sheets/enumeration/google-dorking#13-dorkforge--the-companion-script), a script that builds and explains these for you. + > [!example] Google / GitHub / Shodan dork cheatsheet > | Engine | Dork | What it finds | > |---|---|---| diff --git a/src/content/sheets/pentest-workflow/tool-index.md b/src/content/sheets/pentest-workflow/tool-index.md @@ -35,7 +35,7 @@ source: "vault:Pentest Attack Flow/17 - Tool Index.md" | Stage | Tools | Local attachments | Deep-dive note | | :-- | :-- | :-- | :-- | -| 0 Passive Recon | [crt.sh](https://crt.sh), dig, [Shodan](https://www.shodan.io), google dorks, [gitleaks](https://github.com/gitleaks/gitleaks), [trufflehog](https://github.com/trufflesecurity/trufflehog), [GrayHatWarfare](https://grayhatwarfare.com), [theHarvester](https://github.com/laramies/theHarvester), [subfinder](https://github.com/projectdiscovery/subfinder), [amass](https://github.com/owasp-amass/amass) | — | 2.0 - Cheatsheet - Infrastructure Enumeration Tools · 2.4 - Cheatsheet - Gitleaks · 2.5 - Cheatsheet - TruffleHog · [01 - Stage 00 - Passive External Recon](/sheets/pentest-workflow/passive-external-recon) | +| 0 Passive Recon | [crt.sh](https://crt.sh), dig, [Shodan](https://www.shodan.io), [google dorks](/sheets/enumeration/google-dorking), [gitleaks](https://github.com/gitleaks/gitleaks), [trufflehog](https://github.com/trufflesecurity/trufflehog), [GrayHatWarfare](https://grayhatwarfare.com), [theHarvester](https://github.com/laramies/theHarvester), [subfinder](https://github.com/projectdiscovery/subfinder), [amass](https://github.com/owasp-amass/amass) | — | 2.0 - Cheatsheet - Infrastructure Enumeration Tools · 2.4 - Cheatsheet - Gitleaks · 2.5 - Cheatsheet - TruffleHog · [01 - Stage 00 - Passive External Recon](/sheets/pentest-workflow/passive-external-recon) | | 1 Recon / Host Discovery | [RustScan](https://github.com/bee-san/RustScan), [nmap](https://github.com/nmap/nmap), [masscan](https://github.com/robertdavidgraham/masscan), [naabu](https://github.com/projectdiscovery/naabu), [faketime](https://github.com/wolfcw/libfaketime), ntpdate | — | RustScan_Cheatsheet · Nmap Cheatsheet 2026 · faketime-cheatsheet · [02 - Stage 01 - Recon and Host Discovery](/sheets/pentest-workflow/recon-and-host-discovery) | | 2 Web | [ffuf](https://github.com/ffuf/ffuf), [feroxbuster](https://github.com/epi052/feroxbuster), [gobuster](https://github.com/OJ/gobuster), [wpscan](https://github.com/wpscanteam/wpscan), [nuclei](https://github.com/projectdiscovery/nuclei), [sqlmap](https://github.com/sqlmapproject/sqlmap), [Burp](https://portswigger.net/burp)/[ZAP](https://github.com/zaproxy/zaproxy), [hydra](https://github.com/vanhauser-thc/thc-hydra), [dnSpy](https://github.com/dnSpy/dnSpy), [EyeWitness](https://github.com/RedSiege/EyeWitness), [gowitness](https://github.com/sensepost/gowitness), [ysoserial](https://github.com/frohoff/ysoserial), [ysoserial.net](https://github.com/pwntester/ysoserial.net) | [gowitness_linux_amd64](/downloads/pentest-workflow/gowitness_linux_amd64) ([SHA-256](/downloads/pentest-workflow/gowitness_linux_amd64.sha256) · [GPG signature](/downloads/pentest-workflow/gowitness_linux_amd64.sha256.asc)) [gowitness_windows_amd64.exe](/downloads/pentest-workflow/gowitness_windows_amd64.exe) ([SHA-256](/downloads/pentest-workflow/gowitness_windows_amd64.exe.sha256) · [GPG signature](/downloads/pentest-workflow/gowitness_windows_amd64.exe.sha256.asc)) [ysoserial-all.jar](/downloads/pentest-workflow/ysoserial-all.jar) ([SHA-256](/downloads/pentest-workflow/ysoserial-all.jar.sha256) · [GPG signature](/downloads/pentest-workflow/ysoserial-all.jar.sha256.asc)) [ysoserial.net_v1.36.zip](/downloads/pentest-workflow/ysoserial.net_v1.36.zip) ([SHA-256](/downloads/pentest-workflow/ysoserial.net_v1.36.zip.sha256) · [GPG signature](/downloads/pentest-workflow/ysoserial.net_v1.36.zip.sha256.asc)) [nt-webshell-rosepine.aspx](/downloads/pentest-workflow/nt-webshell-rosepine.aspx) ([SHA-256](/downloads/pentest-workflow/nt-webshell-rosepine.aspx.sha256) · [GPG signature](/downloads/pentest-workflow/nt-webshell-rosepine.aspx.sha256.asc)) [rp-shell.asp](/downloads/pentest-workflow/rp-shell.asp) ([SHA-256](/downloads/pentest-workflow/rp-shell.asp.sha256) · [GPG signature](/downloads/pentest-workflow/rp-shell.asp.sha256.asc)) [rp-shell.jsp](/downloads/pentest-workflow/rp-shell.jsp) ([SHA-256](/downloads/pentest-workflow/rp-shell.jsp.sha256) · [GPG signature](/downloads/pentest-workflow/rp-shell.jsp.sha256.asc)) [rp-shell.php](/downloads/pentest-workflow/rp-shell.php) ([SHA-256](/downloads/pentest-workflow/rp-shell.php.sha256) · [GPG signature](/downloads/pentest-workflow/rp-shell.php.sha256.asc)) | Ffuf-Cheatsheet · WPScan · sqlmap · LFI - Cheat Sheet · 12 - Attacking Thick Client Applications · [03 - Stage 02 - Web Enumeration and Exploitation](/sheets/pentest-workflow/web-enumeration-and-exploitation) | | ⚙ Foothold Toolkit | reverse shells, msfvenom, [Metasploit](https://github.com/rapid7/metasploit-framework), file transfer (certutil / smbserver / curl), [nishang](https://github.com/samratashok/nishang), nc64, [donut](https://github.com/TheWover/donut), [ConPtyShell](https://github.com/antonioCoco/ConPtyShell), [RunasCs](https://github.com/antonioCoco/RunasCs) | [nc64.exe](/downloads/pentest-workflow/nc64.exe) ([SHA-256](/downloads/pentest-workflow/nc64.exe.sha256) · [GPG signature](/downloads/pentest-workflow/nc64.exe.sha256.asc)) [nishang-master.zip](/downloads/pentest-workflow/nishang-master.zip) ([SHA-256](/downloads/pentest-workflow/nishang-master.zip.sha256) · [GPG signature](/downloads/pentest-workflow/nishang-master.zip.sha256.asc)) [donut_v1.1.zip](/downloads/pentest-workflow/donut_v1.1.zip) ([SHA-256](/downloads/pentest-workflow/donut_v1.1.zip.sha256) · [GPG signature](/downloads/pentest-workflow/donut_v1.1.zip.sha256.asc)) | Impacket-Cheatsheet · meterpreter · smbserver.py · [04 - Foothold Toolkit - File Transfers](/sheets/pentest-workflow/foothold-file-transfers) · [05 - Foothold Toolkit - Shells Payloads and Metasploit](/sheets/pentest-workflow/foothold-shells-payloads-metasploit) |