cheatsheet-infrastructure-enumeration-tools-1.md (25821B)
1 --- 2 title: "Cheatsheet - Infrastructure Enumeration Tools 1" 3 description: "curl -s \"https://crt.sh/?q=TARGET.com&output=json\" | jq -r '.[].name_value' | sort -u" 4 category: enumeration 5 tags: ["enumeration"] 6 tools: ["Gitleaks", "TruffleHog"] 7 difficulty: intermediate 8 updated: "2026-08-10" 9 source: "vault:Enumeration/Cheatsheet - Infrastructure Enumeration Tools 1.md" 10 --- 11 # Basic query — good starting point 12 curl -s "https://crt.sh/?q=TARGET.com&output=json" | jq -r '.[].name_value' | sort -u 13 14 # Output 15 *.TARGET.com 16 admin.TARGET.com 17 blog.TARGET.com 18 mail.TARGET.com 19 vpn.TARGET.com 20 www.TARGET.com 21 ``` 22 23 > [!info]+ Command Breakdown 24 > 1. **jq -r '.[].name_value'** — extracts ONLY the `name_value` field from every array entry using `-r` (raw output, no quotes) 25 > 2. This is cleaner than the `grep | cut | awk` chain in the original notes — fewer failure points 26 > 3. **sort -u** — deduplicates; wildcard entries like `*.TARGET.com` are preserved as-is 27 28 --- 29 30 ```bash 31 # Strip wildcards and get clean hostnames only 32 curl -s "https://crt.sh/?q=TARGET.com&output=json" \ 33 | jq -r '.[].name_value' \ 34 | sed 's/\*\.//g' \ 35 | sort -u \ 36 | grep -v "^TARGET.com$" > subdomains.txt 37 38 cat subdomains.txt 39 admin.TARGET.com 40 blog.TARGET.com 41 mail.TARGET.com 42 vpn.TARGET.com 43 www.TARGET.com 44 ``` 45 46 > [!info]+ Command Breakdown 47 > 1. **sed 's/\*\.//g'** — strips the `*.` wildcard prefix to leave a clean hostname 48 > 2. **grep -v "^TARGET.com$"** — removes the bare root domain from the list (you already know it) 49 > 3. **> subdomains.txt** — saves to file for use in the next steps (host loop, Shodan loop) 50 > 4. *This output feeds directly into the `host` bulk resolution loop* 51 52 --- 53 54 ```bash 55 # Query for EXPIRED certs too — these reveal old subdomains that may still be live 56 curl -s "https://crt.sh/?q=%25.TARGET.com&output=json" | jq -r '.[].name_value' | sort -u 57 ``` 58 59 > [!info]+ Command Breakdown 60 > 1. **%25.TARGET.com** — URL-encoded `%` wildcard; matches ALL subdomains ever issued a cert, including expired ones 61 > 2. Expired subdomains are often forgotten by admins — they may still resolve and run old, unpatched services 62 > 3. *Cross-reference this list against your `host` resolution output to see which old subdomains are still live* 63 64 --- 65 66 > [!warning]+ crt.sh Gotchas 67 > 1. Results include **third-party issued certs** — a subdomain listed here is NOT guaranteed to be company-owned infrastructure 68 > 2. Wildcard certs (`*.TARGET.com`) confirm the domain uses wildcard SSL but don't enumerate actual subdomains — use other methods to enumerate what lives under the wildcard 69 > 3. **Rate limiting** — if querying many domains, add `sleep 2` between requests or use the Shodan/Subfinder pipeline instead 70 > 4. Results can lag by hours after a new cert is issued — not real-time 71 72 --- 73 74 ## curl + jq — API Querying and JSON Parsing 75 76 > [!tip]+ jq is More Powerful Than Used in the Notes 77 > The original notes use `jq .` (pretty print only). Here are more useful filters: 78 79 ```bash 80 # Extract only specific fields — issuer + subdomain + expiry 81 curl -s "https://crt.sh/?q=TARGET.com&output=json" \ 82 | jq -r '.[] | [.issuer_name, .name_value, .not_after] | @tsv' 83 84 # Output (tab-separated) 85 C=US, O=Let's Encrypt, CN=R3 mail.TARGET.com 2025-12-01T00:00:00 86 C=US, O=Cloudflare, Inc. www.TARGET.com 2026-01-15T00:00:00 87 ``` 88 89 > [!info]+ Command Breakdown 90 > 1. **'.[] | ...'** — iterates over every element in the array 91 > 2. **[.issuer_name, .name_value, .not_after]** — selects three specific fields into an array 92 > 3. **@tsv** — formats the array as tab-separated values — easy to `cut`, `grep`, or import into a spreadsheet 93 > 4. *The issuer column immediately tells you if a subdomain uses Cloudflare, Let's Encrypt, DigiCert etc. — Cloudflare = WAF/proxy likely in front* 94 95 --- 96 97 ```bash 98 # Count how many unique subdomains exist 99 curl -s "https://crt.sh/?q=TARGET.com&output=json" | jq '[.[].name_value] | unique | length' 100 101 # Output 102 47 103 ``` 104 105 > [!info]+ Command Breakdown 106 > 1. **[.[].name_value]** — collects all name values into a new array 107 > 2. **unique** — deduplicates the array 108 > 3. **length** — returns the count 109 > 4. *Use this as a quick "how big is the attack surface" indicator before diving in* 110 111 --- 112 113 > [!tip]+ curl Best Practices 114 > 1. Always use **-s** (silent) to suppress the progress bar — it pollutes piped output 115 > 2. Add **-A "Mozilla/5.0"** to spoof a browser User-Agent if a service rejects `curl` requests 116 > 3. Use **-o /dev/null -w "%{http_code}"** to test if a URL is reachable without dumping the body 117 > 4. Add **--max-time 10** to prevent curl hanging indefinitely on slow hosts 118 119 --- 120 121 ## dig — DNS Enumeration 122 123 > [!tip]+ Query Specific Record Types — Don't Always Use ANY 124 > Many resolvers block or ignore `dig any` queries. Query record types directly for reliable results: 125 126 ```bash 127 # A records — IPv4 addresses 128 dig A TARGET.com +short 129 130 # AAAA records — IPv6 addresses 131 dig AAAA TARGET.com +short 132 133 # MX records — mail servers 134 dig MX TARGET.com +short 135 136 # NS records — name servers 137 dig NS TARGET.com +short 138 139 # TXT records — the intelligence goldmine 140 dig TXT TARGET.com +short 141 142 # SOA record — primary DNS authority 143 dig SOA TARGET.com +short 144 145 # CNAME — canonical name (reveals CDN, load balancer names) 146 dig CNAME www.TARGET.com +short 147 ``` 148 149 > [!info]+ Command Breakdown 150 > 1. **+short** — suppresses all header/footer output, returns only the answer — much cleaner than default output 151 > 2. Query each record type **separately** — `dig any` often returns incomplete or filtered results from modern resolvers 152 > 3. **CNAME records** are especially useful — they often reveal the CDN or load balancer in use (e.g., `target.cloudflare.net`, `target.azurefd.net`) 153 > 4. *MX pointing to Google = Google Workspace. MX pointing to outlook.com = Microsoft 365. Both have implications for cloud access.* 154 155 --- 156 157 ```bash 158 # Use a specific DNS resolver — bypass internal resolver caching 159 dig TXT TARGET.com @8.8.8.8 +short # Google 160 dig TXT TARGET.com @1.1.1.1 +short # Cloudflare 161 dig TXT TARGET.com @9.9.9.9 +short # Quad9 162 ``` 163 164 > [!info]+ Command Breakdown 165 > 1. **@8.8.8.8** — directs the query to a specific resolver instead of your system's default 166 > 2. Use this when your local resolver returns stale/cached results or when testing from behind a corporate network 167 > 3. Comparing responses between resolvers can reveal **split-horizon DNS** — different answers for internal vs external queries 168 > 4. *Split-horizon DNS is a strong indicator of an internal network — note it for the Gateway layer* 169 170 --- 171 172 ```bash 173 # Attempt a zone transfer — almost always fails externally but worth trying 174 dig AXFR TARGET.com @ns.TARGET.com 175 176 # If it works, output dumps ALL DNS records at once: 177 ; <<>> DiG 9.16.1-Ubuntu <<>> AXFR TARGET.com @ns.TARGET.com 178 TARGET.com. 3600 IN SOA ns1.TARGET.com. hostmaster.TARGET.com. 179 admin.TARGET.com. 3600 IN A 10.10.10.5 180 dev.TARGET.com. 3600 IN A 10.10.10.12 181 vpn.TARGET.com. 3600 IN A 10.10.10.1 182 ``` 183 184 > [!info]+ Command Breakdown 185 > 1. **AXFR** — DNS zone transfer request; asks the name server to send ALL records for the zone 186 > 2. Requires knowing the authoritative NS server — get this from `dig NS TARGET.com +short` first 187 > 3. Modern DNS servers reject AXFR from untrusted IPs, but misconfigured or older servers may allow it 188 > 4. *A successful zone transfer is a critical finding — it immediately hands you the entire internal DNS map* 189 190 > [!warning]+ dig vs host vs nslookup 191 > | Tool | Best For | Avoid When | 192 > |---|---|---| 193 > | `dig` | Full record queries, scripting, verbose output | You just need a quick IP lookup | 194 > | `host` | Fast bulk lookups, simple A record resolution | You need specific record types or JSON output | 195 > | `nslookup` | Windows environments | Scripting — its output format is inconsistent | 196 197 --- 198 199 ## host — Bulk Subdomain Resolution 200 201 > [!tip]+ Going Further Than the Notes 202 203 ```bash 204 # Reverse DNS lookup — IP to hostname 205 host 10.129.24.93 206 93.24.129.10.in-addr.arpa domain name pointer blog.TARGET.com. 207 ``` 208 209 > [!info]+ Command Breakdown 210 > 1. Pass an **IP address** to `host` instead of a hostname for reverse lookup 211 > 2. Reveals hostnames you may not have found in forward DNS enumeration 212 > 3. *Run this on every IP returned by Shodan — you may discover additional virtual hosts pointing to the same IP* 213 214 --- 215 216 ```bash 217 # Query a specific record type with host 218 host -t MX TARGET.com 219 host -t TXT TARGET.com 220 host -t NS TARGET.com 221 ``` 222 223 > [!info]+ Command Breakdown 224 > 1. **-t [TYPE]** — specifies the record type to query 225 > 2. Output is simpler than `dig` — useful for quick checks but less detail 226 > 3. *Use `dig` for full output with TTL and authority sections; use `host -t` for fast piped workflows* 227 228 --- 229 230 ```bash 231 # Better bulk resolution loop — saves both hostname and IP, skips failed lookups 232 while read subdomain; do 233 result=$(host "$subdomain" 2>/dev/null | grep "has address" | awk '{print $1, $4}') 234 [ -n "$result" ] && echo "$result" 235 done < subdomains.txt | tee resolved.txt 236 237 # Output 238 blog.TARGET.com 10.129.24.93 239 mail.TARGET.com 10.129.127.22 240 www.TARGET.com 10.129.127.33 241 ``` 242 243 > [!info]+ Command Breakdown 244 > 1. **while read subdomain** — cleaner than `for i in $(cat file)` — handles spaces in lines safely 245 > 2. **2>/dev/null** — suppresses error output for subdomains that don't resolve 246 > 3. **[ -n "$result" ]** — only prints if the result is non-empty (skips dead subdomains silently) 247 > 4. **tee resolved.txt** — prints to terminal AND saves to file simultaneously 248 > 5. *`resolved.txt` becomes your definitive list of live, company-hosted targets for active testing* 249 250 --- 251 252 ## Shodan — Passive IP and Service Intelligence 253 254 > [!important]+ Setup First — Easy to Skip 255 ```bash 256 # Install the Shodan CLI 257 pip3 install shodan 258 259 # Initialise with your API key (free account works for basic queries) 260 shodan init YOUR_API_KEY_HERE 261 262 # Verify it works 263 shodan info 264 # Output: 265 # Query credits available: 100 266 # Scan credits available: 0 267 ``` 268 269 > [!info]+ Command Breakdown 270 > 1. **shodan init** — stores your API key locally so you don't have to pass it every command 271 > 2. Free accounts get 100 query credits — enough for a standard engagement's passive recon 272 > 3. *Get your API key at https://account.shodan.io/ — just needs a free registration* 273 274 --- 275 276 ```bash 277 # The basic host lookup from the notes 278 shodan host 10.129.24.93 279 280 # Better: search by organisation name — finds ALL IPs registered to the company 281 shodan search --fields ip_str,port,org,hostnames "org:\"InlaneFreight\"" 282 283 # Output 284 10.129.24.93 80 InlaneFreight blog.inlanefreight.com 285 10.129.27.33 443 InlaneFreight www.inlanefreight.com 286 10.129.127.22 25 InlaneFreight matomo.inlanefreight.com 287 ``` 288 289 > [!info]+ Command Breakdown 290 > 1. **"org:\"InlaneFreight\""** — Shodan search filter for organisation name; quotes inside the string are escaped 291 > 2. **--fields ip_str,port,org,hostnames** — limits output to only the useful columns 292 > 3. *This finds IPs you may have missed entirely in DNS enumeration — Shodan indexes based on what it scans, not what DNS says* 293 294 --- 295 296 ```bash 297 # Find all subdomains Shodan has seen for a domain 298 shodan search --fields ip_str,port,hostnames "hostname:TARGET.com" 299 300 # Find only SSL-enabled services — check for weak ciphers 301 shodan search "hostname:TARGET.com ssl.version:TLSv1" 302 303 # Find specific open ports across the org 304 shodan search "org:\"InlaneFreight\" port:22" 305 shodan search "org:\"InlaneFreight\" port:3389" # RDP — always worth noting 306 shodan search "org:\"InlaneFreight\" port:445" # SMB — goldmine 307 308 # Find services with default/vendor credentials (Shodan tags these) 309 shodan search "org:\"InlaneFreight\" has_screenshot:true" 310 ``` 311 312 > [!info]+ Command Breakdown 313 > 1. **hostname:TARGET.com** — finds all IPs where Shodan has observed the domain in the SSL cert or reverse DNS 314 > 2. **ssl.version:TLSv1** — filters for hosts still running deprecated TLS 1.0 — a vulnerability 315 > 3. **port:3389** / **port:445** — RDP and SMB exposed to the internet are high-priority findings 316 > 4. **has_screenshot:true** — Shodan captures screenshots of HTTP services — you can see login panels, admin interfaces, and dashboards without sending a single packet to the target 317 318 --- 319 320 > [!tip]+ Shodan Filters Quick Reference 321 322 | Filter | Example | What It Finds | 323 |---|---|---| 324 | `org:` | `org:"Target Corp"` | All IPs registered to that organisation | 325 | `hostname:` | `hostname:target.com` | IPs with that hostname in cert/DNS | 326 | `port:` | `port:22` | Services on a specific port | 327 | `country:` | `country:GB` | Restrict by country | 328 | `ssl.version:` | `ssl.version:TLSv1` | Weak SSL/TLS versions | 329 | `product:` | `product:Apache` | Specific software | 330 | `os:` | `os:"Windows Server 2016"` | Specific OS versions | 331 | `has_screenshot:` | `has_screenshot:true` | Services with captured screenshots | 332 | `http.title:` | `http.title:"Login"` | Pages with specific HTML titles | 333 334 --- 335 336 ## Google Dorks — Cloud and File Discovery 337 338 > [!tip]+ Beyond the Basic inurl/intext Combo 339 > The notes cover `inurl:` and `intext:`. Here is the full operator toolkit: 340 341 ``` 342 # Find exposed files by type on the company's domain 343 site:TARGET.com filetype:pdf 344 site:TARGET.com filetype:xlsx 345 site:TARGET.com filetype:docx 346 site:TARGET.com filetype:env 347 site:TARGET.com filetype:sql 348 site:TARGET.com filetype:log 349 ``` 350 351 > [!info]+ Command Breakdown 352 > 1. **site:TARGET.com** — restricts ALL results to the specified domain and its subdomains 353 > 2. **filetype:** — filters by file extension — find documents, spreadsheets, SQL dumps, log files 354 > 3. `.env` files indexed by Google are an instant win — they almost always contain secrets 355 > 4. `.sql` files may contain database dumps with credentials and PII 356 357 --- 358 359 ``` 360 # Find login panels and admin interfaces 361 site:TARGET.com intitle:"login" 362 site:TARGET.com intitle:"admin" 363 site:TARGET.com inurl:"/admin" 364 site:TARGET.com inurl:"/wp-admin" 365 site:TARGET.com inurl:"/phpmyadmin" 366 site:TARGET.com inurl:"/dashboard" 367 ``` 368 369 > [!info]+ Command Breakdown 370 > 1. **intitle:** — searches the HTML `<title>` tag of pages — login panels almost always say "Login" in their title 371 > 2. **inurl:** — searches the URL path — admin panels follow predictable URL patterns 372 > 3. *Combine `site:` with `intitle:` for targeted results with zero noise* 373 374 --- 375 376 ``` 377 # Cloud storage discovery — go beyond the notes 378 intext:"TARGET" inurl:amazonaws.com 379 intext:"TARGET" inurl:blob.core.windows.net 380 intext:"TARGET" inurl:storage.googleapis.com 381 intext:"TARGET" inurl:s3.amazonaws.com 382 intext:"TARGET" inurl:digitaloceanspaces.com 383 384 # Find cached/old versions of pages 385 cache:TARGET.com/admin 386 387 # Find subdomains not in DNS — Google has indexed them 388 site:*.TARGET.com -site:www.TARGET.com 389 ``` 390 391 > [!info]+ Command Breakdown 392 > 1. **storage.googleapis.com** and **digitaloceanspaces.com** — additional cloud storage providers beyond AWS/Azure that are often forgotten 393 > 2. **cache:TARGET.com/page** — Google's cached version of a page — may show content from before a page was taken down or secured 394 > 3. **site:\*.TARGET.com -site:www.TARGET.com** — the `-` operator excludes a site; this surfaces all indexed subdomains except `www` — a fast way to discover what Google has crawled 395 396 --- 397 398 > [!tip]+ Google Dork Pro Tips 399 > 1. Use **"quotes"** around exact phrases — `"internal use only"` finds documents accidentally published 400 > 2. Combine multiple operators: `site:TARGET.com filetype:pdf intitle:"confidential"` 401 > 3. Use the **`-`** operator to subtract noise: `site:TARGET.com -site:blog.TARGET.com` 402 > 4. Google limits dork results — if you hit ~30 results, rephrase or use different operators for fresh results 403 > 5. **Never use Google Dorks logged into your Google account** during an engagement — your search history is tied to your identity 404 405 --- 406 407 ## GrayHatWarfare — Cloud Bucket Enumeration 408 409 > [!tip]+ Effective Search Strategy 410 > The notes say "search by company name" — here is a systematic approach: 411 412 ``` 413 Search terms to try (in order): 414 1. Full company name: "InlaneFreight" 415 2. Common abbreviation: "ILF" 416 3. Domain without TLD: "inlanefreight" 417 4. Product/brand names: [any product names found on the website] 418 5. Internal codenames: [any codenames found in job postings or GitHub] 419 ``` 420 421 > [!info]+ File Types to Prioritise 422 423 | Priority | File Type | Why | 424 |---|---|---| 425 | **Critical** | `id_rsa`, `id_ecdsa`, `.pem`, `.ppk` | SSH/TLS private keys — direct access | 426 | **Critical** | `.env`, `*.env.production` | API keys, DB passwords, secrets | 427 | **High** | `config.json`, `settings.py`, `web.config` | Hardcoded credentials, internal IPs | 428 | **High** | `*.sql`, `*.db`, `*.sqlite` | Database dumps — credentials + data | 429 | **Medium** | `*.xlsx`, `*.csv` | May contain employee lists, IP lists, passwords | 430 | **Medium** | `*.pdf` | Internal documentation, network diagrams | 431 | **Low** | `*.log` | May contain tokens, paths, usernames in entries | 432 433 --- 434 435 > [!warning]+ GrayHatWarfare Limitations 436 > 1. Only indexes **publicly accessible** buckets — password-protected or private buckets won't appear 437 > 2. Its index is not real-time — newly exposed buckets may take days to appear 438 > 3. Files listed may have since been removed — always verify before reporting 439 > 4. **Do NOT download files without explicit written scope permission** — accessing unauthenticated cloud storage may still carry legal risk depending on jurisdiction 440 441 --- 442 443 ## domain.glass — Quick Infrastructure Snapshot 444 445 > [!tip]+ What to Actually Look At 446 > The notes mention it exists. Here is what to focus on when you open it: 447 448 ``` 449 URL: https://domain.glass/TARGET.com 450 ``` 451 452 > [!info]+ Sections That Matter 453 454 | Section | What to Extract | 455 |---|---| 456 | **IP Information** | Hosting provider, ASN, IP range — tells you who hosts the infrastructure | 457 | **SSL Certificate** | Issuer (Cloudflare? Let's Encrypt? Internal CA?) and listed SANs (extra subdomains) | 458 | **Cloudflare Status** | "Safe" = Cloudflare is proxying — real IP is hidden; direct IP scanning will hit Cloudflare, not the origin | 459 | **DNS Records** | Cross-reference against your `dig` output — domain.glass sometimes catches records `dig any` misses | 460 | **Social Media Links** | Auto-detected official accounts — cross-reference with your LinkedIn/staff OSINT | 461 462 > [!warning]+ Cloudflare Proxy Implication 463 > 1. If domain.glass shows Cloudflare as "Safe" — the A record IP is a **Cloudflare IP, not the origin server** 464 > 2. Active scanning against that IP hits Cloudflare's infrastructure — not the target 465 > 3. Find the real IP via: historical DNS records (SecurityTrails), SSL cert transparency, Shodan `ssl.cert.subject.cn:TARGET.com`, or email headers (MX trace) 466 467 --- 468 469 ## LinkedIn — Staff and Tech Stack OSINT 470 471 > [!tip]+ Search Syntax That Actually Works 472 > LinkedIn's search is intentionally limited for free accounts. Here is how to work around it: 473 474 ``` 475 # Use Google to search LinkedIn profiles — more powerful than LinkedIn's own search 476 site:linkedin.com/in "TARGET company name" "software engineer" 477 site:linkedin.com/in "TARGET company name" "security engineer" 478 site:linkedin.com/in "TARGET company name" "devops" 479 site:linkedin.com/in "TARGET company name" "sysadmin" 480 site:linkedin.com/in "TARGET company name" "AWS" OR "Azure" OR "GCP" 481 ``` 482 483 > [!info]+ Command Breakdown 484 > 1. **site:linkedin.com/in** — restricts Google results to LinkedIn profile pages only 485 > 2. Combine the company name with job titles to surface the most relevant staff 486 > 3. Add technology keywords (`"AWS"`, `"Kubernetes"`, `"Splunk"`) to find staff who list those skills 487 > 4. *Google's index of LinkedIn is deeper than LinkedIn's own search — especially useful without a Premium account* 488 489 --- 490 491 > [!tip]+ What to Record from Each Profile 492 493 | Profile Section | What to Note | 494 |---|---| 495 | **Current Role + Company** | Confirms employment — verify you have the right person | 496 | **Tech Stack in "About"** | Programming languages, cloud platforms, tools in active use | 497 | **GitHub / Portfolio Links** | Jump to code search immediately — these are the highest-value leads | 498 | **Career History** | Technologies used at each role — older systems may still be in use | 499 | **Certifications** | AWS/Azure certified = likely manages cloud; OSCP/CEH = security awareness is higher | 500 | **Followed Companies** | May indicate vendors they use or are evaluating | 501 | **Activity / Posts** | Recent posts about specific tools = those tools are in active use RIGHT NOW | 502 503 --- 504 505 > [!tip]+ Priority Targets on LinkedIn 506 > 1. **IT/Infrastructure admins** — they configure the systems you're targeting 507 > 2. **DevOps / SRE engineers** — they wrote the pipelines that deploy to cloud 508 > 3. **Security engineers** — their skills tell you what defences exist (SIEM? EDR? WAF?) 509 > 4. **Former employees** — may have older, potentially still-valid credentials; less cautious about what they share post-employment 510 > 5. **Junior developers** — more likely to have public GitHub repos with company-adjacent code 511 512 --- 513 514 ## GitHub — Code and Secret Hunting 515 516 > [!tip]+ GitHub Search Syntax Goes Far Beyond the Notes 517 > Most recon stops at "browse the repo". GitHub has a powerful search API: 518 519 ``` 520 # Search for company name in ALL public code 521 org:TARGET-org-name # All repos under a specific GitHub org 522 user:firstname-lastname TARGET # Repos by a specific employee mentioning the company 523 524 # Search for secrets by filename 525 filename:.env TARGET 526 filename:config.py password 527 filename:settings.py SECRET_KEY 528 filename:application.properties datasource 529 530 # Search for secrets by content 531 "TARGET.com" password 532 "TARGET.com" api_key 533 "TARGET.com" BEGIN RSA PRIVATE KEY 534 "TARGET.com" token 535 536 # Search for internal infrastructure hints 537 "TARGET.com" internal 538 "TARGET.com" 192.168. 539 "TARGET.com" 10.0. 540 "TARGET.com" staging 541 "TARGET.com" production 542 ``` 543 544 > [!info]+ Command Breakdown 545 > 1. **org:** — restricts search to all repositories under a GitHub organisation (the company may have an official org) 546 > 2. **filename:** — searches ONLY in files with that specific name — highly targeted for common secret files 547 > 3. **"BEGIN RSA PRIVATE KEY"** — literal string that starts every RSA private key — if this appears in a repo, it's a critical finding 548 > 4. Internal IP ranges (`192.168.`, `10.0.`) in public repos reveal internal network addressing 549 550 --- 551 552 ```bash 553 # Check commit history for deleted secrets — CLI approach 554 git clone https://github.com/target-org/repo-name 555 cd repo-name 556 557 # Search ALL commits for the word "password" 558 git log --all -p | grep -i "password" | head -50 559 560 # Search for a specific string across all branches and commits 561 git log --all --oneline | awk '{print $1}' | xargs -I{} git grep -l "SECRET_KEY" {} 562 ``` 563 564 > [!info]+ Command Breakdown 565 > 1. **git log --all -p** — shows every commit across all branches WITH the full diff (added/removed lines) 566 > 2. **grep -i "password"** — case-insensitive search through all commit diffs 567 > 3. Even if a secret was removed in a later commit, `git log -p` shows the line prefixed with `+` (added) and `-` (removed) — deleted secrets are still readable in the diff 568 > 4. *This is how most credential leaks persist — the file is "cleaned up" but the history is never purged* 569 570 --- 571 572 > [!tip]+ Automated Secret Scanning Tools (Beyond Manual Search) 573 574 | Tool | Command | What It Does | 575 |---|---|---| 576 | [trufflehog](https://github.com/trufflesecurity/trufflehog) | `trufflehog github --org=TARGET-org` | Scans all org repos + full history for 700+ secret patterns | 577 | [gitleaks](https://github.com/gitleaks/gitleaks) | `gitleaks detect --source=./repo` | Scans local repo for secrets using regex rules | 578 | [gitrob](https://github.com/michenriksen/gitrob) | `gitrob TARGET-org` | Maps org members' repos and scans for sensitive files | 579 580 > [!warning]+ GitHub OSINT Rules 581 > 1. Only search **public** repositories — accessing private repos without authorisation is illegal 582 > 2. **Do not clone or download** any repo that isn't clearly in scope — downloading may constitute unauthorised access in some jurisdictions 583 > 3. GitHub may rate-limit unauthenticated searches — authenticate with a **throwaway account** if needed, never your real account 584 > 4. Findings from GitHub (leaked keys, hardcoded passwords) must be **reported immediately** in real engagements — they are often actively exploitable 585 586 --- 587 588 ## References 589 590 1. [HTB Academy - Footprinting Module](https://academy.hackthebox.com/module/details/112) 591 2. [crt.sh - Certificate Transparency](https://crt.sh/) 592 3. [Shodan CLI Documentation](https://cli.shodan.io/) 593 4. [Shodan Search Filters Reference](https://www.shodan.io/search/filters) 594 5. [dig Man Page](https://linux.die.net/man/1/dig) 595 6. [jq Manual](https://stedolan.github.io/jq/manual/) 596 7. [Google Advanced Search Operators](https://support.google.com/websearch/answer/2466433) 597 8. [GrayHatWarfare](https://buckets.grayhatwarfare.com/) 598 9. [domain.glass](https://domain.glass/) 599 10. [TruffleHog - Secret Scanner](https://github.com/trufflesecurity/trufflehog) 600 11. [Gitleaks](https://github.com/gitleaks/gitleaks) 601 12. [HackTricks - DNS Enumeration](https://book.hacktricks.xyz/network-services-pentesting/pentesting-dns) 602 13. [HackTricks - External Recon Methodology](https://book.hacktricks.xyz/generic-methodologies-and-resources/external-recon-methodology) 603 14. [MITRE ATT&CK - Search Open Technical Databases (T1596)](https://attack.mitre.org/techniques/T1596/) 604 15. [MITRE ATT&CK - Search Open Websites/Domains (T1593)](https://attack.mitre.org/techniques/T1593/) 605 606 --- 607 608 #HTB #Footprinting #OSINT #Cheatsheet #DNS #Shodan #GoogleDorking #CertificateTransparency #GrayHatWarfare #GitHub #LinkedIn #SubdomainEnumeration #CloudStorage #PassiveRecon