Crafting Boolean Search Dorks
Markdown--- name: crafting-boolean-search-dorks description: Constructs precise Boolean search strings and Google Dorking queries to bypass AI-generated summaries and retrieve literal, token-matched results from search engines. Use when a user needs to find specific files, exposed data, pre-AI/human-authored content, or wants to force verbatim/non-AI search behavior instead of conversational semantic search. --- # Crafting Boolean Search Dorks
To find PDF reports on a specific domain, published before generative AI content flooded the web, and strip out AI Overviews:
site:example.com filetype:pdf intitle:"annual report" -ai before:2022-11-30
Append &udm=14 to the resulting Google URL to force the classic text-only, link-first layout (no AI Overview, no auto-expansion).
Progress:
- Step 1: Clarify the retrieval goal (exact file, exposed data, domain-restricted content, or historical/pre-AI content)
- Step 2: Select core Boolean logic (AND / OR / NOT / grouping)
- Step 3: Add relevant Dorking operators (site:, filetype:, inurl:, intitle:, intext:, AROUND(N))
- Step 4: Apply AI-bypass layer if needed (verbatim mode, -ai, udm=14, tbs=li:1)
- Step 5: Apply temporal constraints if isolating pre-AI/human content (before:/after:)
- Step 6: Choose platform (Google verbatim, DuckDuckGo No-AI, Mojeek, SymbolHound) based on need
- Step 7: Validate query syntax (no spaces around operators, quotes around exact phrases, parentheses balanced)
Step-by-step details
-
Define intent. Is this OSINT/recon (dorking), a narrow document lookup, or an attempt to avoid synthesized/AI answers? This determines which operator family to use.
-
Build the Boolean skeleton first.
AND/&— force intersection (often implicit; use explicit AND for clarity in complex strings)OR/|— broaden to synonyms/alternatives — always wrap in parentheses when mixed with AND-term— exclude (no space between-and the term)()— group sub-expressions to control evaluation order
-
Layer in Dorking operators (combine freely, space-separated):
Operator Purpose site:restrict to domain/TLD filetype:/ext:restrict to file extension inurl:string must appear in URL path intitle:string must appear in <title>intext:string must appear in body text AROUND(N)two terms within N words of each other Note:
related:andcache:are deprecated/removed — do not rely on them. -
Force literal/verbatim matching when the engine is auto-correcting, swapping synonyms, or injecting AI summaries:
- Use Google's Verbatim mode (Tools → All Results → Verbatim)
- Append
&udm=14to the URL (text-only web results, no AI Overview) - Append
tbs=li:1(enforces verbatim) - Add
-aias an exclusion token to suppress AI/overview content
-
Apply temporal dorking to isolate pre-generative-AI content:
before:2022-11-30— pre-ChatGPT launch cutoff (most precise)before:2023— clean annual cutoffbefore:2019— strict historical window, pre-transformer-era content- URL-layer equivalent:
&tbs=cdr:1,cd_max=11/29/2022
-
Switch platforms when Google dorking is throttled or summarized away:
- DuckDuckGo No-AI — strips conversational widgets
- Mojeek — independent index, literal token matching, no semantic synthesis
- SymbolHound — searches special characters/symbols that major engines ignore (useful for code, error messages, syntax)
Example 1: Input: "Find exposed admin login pages on .edu domains, excluding AI-related noise." Output:
inurl:admin intitle:"login" site:*.edu -ai
Example 2: Input: "Locate Word resumes mentioning 'machine learning' but not 'internship', published before ChatGPT existed." Output:
filetype:docx intext:"machine learning" -internship before:2022-11-30
Example 3: Input: "Search for pages about 'python OR javascript' tutorials, but not for beginners, and force a non-AI Google results page." Output:
(python OR javascript) tutorial -beginners
Append &udm=14 to the search URL to render classic link-only results.
Example 4: Input: "Find strict pre-2019 forum discussions about 'cold fusion' within 3 words of 'experiment', avoiding modern AI-generated rehashes." Output:
"cold fusion" AROUND(3) experiment before:2019
Run via Mojeek or DuckDuckGo No-AI for additional AI-content filtering.
- Always quote multi-word exact phrases:
"index of"notindex of. - No space after a prefix operator:
filetype:pdfnotfiletype: pdf; same for-exclude. - Wrap OR clauses in parentheses when combined with AND terms to avoid ambiguous precedence.
- Default to
before:2022-11-30when the goal is strictly "pre-AI content" — it's the most defensible, citable cutoff. - Combine
-ai+&udm=14+ Verbatim mode together for maximum suppression of synthesized results; one alone is often insufficient. - When a dork chain gets ignored or "softened" by the engine, switch to Mojeek or DuckDuckGo No-AI rather than fighting Google's query reinterpretation.
- For symbols/code snippets (e.g.,
C++,$variable, regex), use SymbolHound instead of mainstream engines, which strip special characters.
- Don't rely on
related:orcache:— both are deprecated/dead on Google. - Don't assume Boolean operators are absolute on modern engines by default — they're often treated as soft ranking signals unless Verbatim mode or
udm=14is explicitly forced. - Don't put a space between a negative sign and the excluded term (
- termfails;-termworks). - Don't mix OR and AND without parentheses — results in unpredictable precedence.
- Don't expect
before:/after:to be pixel-perfect; they filter by Google's inferred publish/crawl date, which can be inaccurate for undated or republished pages. - Don't forget that excluding
-aifilters the literal token "ai," which can over-exclude legitimate pages (e.g., about "air" compounds) — test and refine.