AI Skill Report Card

Crafting Boolean Search Dorks

A-87·Oct 2, 2026·Source: Extension-selection
Markdown
--- name: crafting-boolean-search-dorks description: Constructs precise Boolean search strings and Google Dorking queries to bypass AI-generated summaries and retrieve literal, token-matched results from search engines. Use when a user needs to find specific files, exposed data, pre-AI/human-authored content, or wants to force verbatim/non-AI search behavior instead of conversational semantic search. --- # Crafting Boolean Search Dorks
14 / 15

To find PDF reports on a specific domain, published before generative AI content flooded the web, and strip out AI Overviews:

site:example.com filetype:pdf intitle:"annual report" -ai before:2022-11-30

Append &udm=14 to the resulting Google URL to force the classic text-only, link-first layout (no AI Overview, no auto-expansion).

Recommendation▾
Add an example of combining multiple AI-bypass techniques simultaneously to show full-stack suppression in action
14 / 15

Progress:

  • Step 1: Clarify the retrieval goal (exact file, exposed data, domain-restricted content, or historical/pre-AI content)
  • Step 2: Select core Boolean logic (AND / OR / NOT / grouping)
  • Step 3: Add relevant Dorking operators (site:, filetype:, inurl:, intitle:, intext:, AROUND(N))
  • Step 4: Apply AI-bypass layer if needed (verbatim mode, -ai, udm=14, tbs=li:1)
  • Step 5: Apply temporal constraints if isolating pre-AI/human content (before:/after:)
  • Step 6: Choose platform (Google verbatim, DuckDuckGo No-AI, Mojeek, SymbolHound) based on need
  • Step 7: Validate query syntax (no spaces around operators, quotes around exact phrases, parentheses balanced)

Step-by-step details

  1. Define intent. Is this OSINT/recon (dorking), a narrow document lookup, or an attempt to avoid synthesized/AI answers? This determines which operator family to use.

  2. Build the Boolean skeleton first.

    • AND / & — force intersection (often implicit; use explicit AND for clarity in complex strings)
    • OR / | — broaden to synonyms/alternatives — always wrap in parentheses when mixed with AND
    • -term — exclude (no space between - and the term)
    • () — group sub-expressions to control evaluation order
  3. Layer in Dorking operators (combine freely, space-separated):

    OperatorPurpose
    site:restrict to domain/TLD
    filetype: / ext:restrict to file extension
    inurl:string must appear in URL path
    intitle:string must appear in <title>
    intext:string must appear in body text
    AROUND(N)two terms within N words of each other

    Note: related: and cache: are deprecated/removed — do not rely on them.

  4. Force literal/verbatim matching when the engine is auto-correcting, swapping synonyms, or injecting AI summaries:

    • Use Google's Verbatim mode (Tools → All Results → Verbatim)
    • Append &udm=14 to the URL (text-only web results, no AI Overview)
    • Append tbs=li:1 (enforces verbatim)
    • Add -ai as an exclusion token to suppress AI/overview content
  5. Apply temporal dorking to isolate pre-generative-AI content:

    • before:2022-11-30 — pre-ChatGPT launch cutoff (most precise)
    • before:2023 — clean annual cutoff
    • before:2019 — strict historical window, pre-transformer-era content
    • URL-layer equivalent: &tbs=cdr:1,cd_max=11/29/2022
  6. Switch platforms when Google dorking is throttled or summarized away:

    • DuckDuckGo No-AI — strips conversational widgets
    • Mojeek — independent index, literal token matching, no semantic synthesis
    • SymbolHound — searches special characters/symbols that major engines ignore (useful for code, error messages, syntax)
Recommendation▾
Include a brief note on legal/ethical boundaries of dorking since it touches OSINT/exposed-data use cases
18 / 20

Example 1: Input: "Find exposed admin login pages on .edu domains, excluding AI-related noise." Output:

inurl:admin intitle:"login" site:*.edu -ai

Example 2: Input: "Locate Word resumes mentioning 'machine learning' but not 'internship', published before ChatGPT existed." Output:

filetype:docx intext:"machine learning" -internship before:2022-11-30

Example 3: Input: "Search for pages about 'python OR javascript' tutorials, but not for beginners, and force a non-AI Google results page." Output:

(python OR javascript) tutorial -beginners

Append &udm=14 to the search URL to render classic link-only results.

Example 4: Input: "Find strict pre-2019 forum discussions about 'cold fusion' within 3 words of 'experiment', avoiding modern AI-generated rehashes." Output:

"cold fusion" AROUND(3) experiment before:2019

Run via Mojeek or DuckDuckGo No-AI for additional AI-content filtering.

Recommendation▾
Consider a troubleshooting table mapping symptoms (e.g., 'results still show AI Overview') to specific fixes
  • Always quote multi-word exact phrases: "index of" not index of.
  • No space after a prefix operator: filetype:pdf not filetype: pdf; same for -exclude.
  • Wrap OR clauses in parentheses when combined with AND terms to avoid ambiguous precedence.
  • Default to before:2022-11-30 when the goal is strictly "pre-AI content" — it's the most defensible, citable cutoff.
  • Combine -ai + &udm=14 + Verbatim mode together for maximum suppression of synthesized results; one alone is often insufficient.
  • When a dork chain gets ignored or "softened" by the engine, switch to Mojeek or DuckDuckGo No-AI rather than fighting Google's query reinterpretation.
  • For symbols/code snippets (e.g., C++, $variable, regex), use SymbolHound instead of mainstream engines, which strip special characters.
  • Don't rely on related: or cache: — both are deprecated/dead on Google.
  • Don't assume Boolean operators are absolute on modern engines by default — they're often treated as soft ranking signals unless Verbatim mode or udm=14 is explicitly forced.
  • Don't put a space between a negative sign and the excluded term (- term fails; -term works).
  • Don't mix OR and AND without parentheses — results in unpredictable precedence.
  • Don't expect before:/after: to be pixel-perfect; they filter by Google's inferred publish/crawl date, which can be inaccurate for undated or republished pages.
  • Don't forget that excluding -ai filters the literal token "ai," which can over-exclude legitimate pages (e.g., about "air" compounds) — test and refine.
0
Grade A-AI Skill Framework
Scorecard
Criteria Breakdown
Quick Start
14/15
Workflow
14/15
Examples
18/20
Completeness
18/20
Format
15/15
Conciseness
14/15