Hunting AI Arbitrage Loops
Markdown--- name: hunting-ai-arbitrage-loops description: Researches and documents zero-cost AI Data Arbitrage and Generative Engine Optimization (GEO) workflows by scanning for underground playbooks, unpatched scraping workarounds, and free-tool monetization pipelines. Use when asked to find growth-hacking loops, GEO/LLMO tactics, free data ingestion-to-publish pipelines, or "nuggets" combining scraping, structuring, and monetizing data via free infrastructure (GitHub Pages, Vercel, MCP servers, free LLM tiers). ---
Given a research request for AI arbitrage / GEO loops, immediately structure findings into the 5-stage pipeline (Buy → Flip → Publish → Profit → Track), not generic marketing advice. Every finding must include three components:
- The Tactical Play (one-sentence loop description)
- The Tech Stack (100% free tools only)
- The Exact Steps (runnable script/prompt/config)
If no live web search is available, construct the most plausible, technically sound version of the playbook from known tool capabilities and documented platform behaviors (e.g., known crawler user-agents, known free-tier API limits, documented GitHub Pages/Vercel behavior). Never fabricate specific URLs, named blog posts, or named authors — describe the mechanism, not a fake citation.
Progress:
- Clarify monetization route and ingestion target if ambiguous (ask once, then proceed with best default)
- Map the "Buy" phase: identify free ingestion mechanism
- Map the "Flip" phase: identify free value-add/formatting mechanism
- Map the "Publish" phase: identify zero-cost hosting mechanism that exploits LLM crawler behavior
- Map the "Profit" phase: identify monetization hook
- Map the "Track" phase: identify free measurement tool
- Compile into bucketed output with tactical play / stack / steps for each
- Flag legal/ToS risk level for each tactic (public record vs. scraping gated content)
Step 1 — Scope the request. If the user hasn't specified a monetization route (affiliate content vs. B2B lead-gen vs. MCP micro-fees vs. dataset licensing), pick the most generalizable default (local public-records lead-gen, since it requires no audience and no product) and note that other routes are swappable.
Step 2 — Build the Buy phase. Default nugget: target datasets that are (a) public record, (b) fragmented across municipal PDFs/HTML tables, (c) not yet aggregated. Use free scraping tools: Python requests + BeautifulSoup, Playwright (free, local), Google Colab for compute, and free-tier Firecrawl/Jina Reader (r.jina.ai/<url>) for clean markdown extraction without a scraping stack. Note the specific unpatched workaround: prefixing any URL with https://r.jina.ai/ returns LLM-ready markdown for free, bypassing the need to build a custom parser.
Step 3 — Build the Flip phase. Use free-tier LLM APIs (Gemini free tier, Groq free tier, or local Ollama models) to: (a) restructure raw scrape into clean JSON, (b) inject "Information Gain" sentences (facts/stats not found in top 10 SERP competitors — feed competitor text in-context and prompt for delta), (c) auto-generate Schema.org JSON-LD blocks matching the entity type (LocalBusiness, GovernmentOrganization, Product).
Step 4 — Build the Publish phase. Default: push structured .md/.json files to a public GitHub repo (GitHub Pages for HTML rendering is optional — raw .md/.json files in a public repo are directly crawlable by OAI-SearchBot, PerplexityBot, ClaudeBot without needing Google indexing at all). Alternative: Vercel/Netlify free tier for a thin static site wrapping the JSON as human-readable pages, with robots.txt deliberately allowing AI crawlers.
Step 5 — Build the Profit phase. Default hooks: (a) embed a free-tier form (Tally, Google Forms) as the CTA on each generated page for lead-gen, (b) wrap the JSON dataset as an MCP server (clone modelcontextprotocol/servers template, point it at the JSON/SQLite file) and monetize via usage-based access for Cursor/Cline users, (c) affiliate links injected contextually in the "Flip" formatting step.
Step 6 — Build the Track phase. Use free/manual GEO tracking: manually query ChatGPT/Perplexity/Gemini with target queries and log whether your domain/dataset is cited (build a simple spreadsheet tracker); use free tiers of LLMrefs, Frase GEO checker, or HubSpot AI Search Grader; alternatively scrape your own site's server logs (Vercel/Netlify logs are free) for hits from OAI-SearchBot, PerplexityBot, ClaudeBot, GPTBot user-agents as a free, direct SoV proxy.
Step 7 — Compile output using the bucket format in Examples below. Always end by offering the two narrowing questions (monetization route, or ingestion script deep-dive) — but only as a follow-up offer, never blocking the initial output.
Example 1: Input: "Find me a zero-cost loop for turning municipal zoning data into a monetizable asset."
Output:
BUCKET: Public Record → Structured Dataset → MCP Monetization
1. Tactical Play: Scrape unstructured municipal zoning PDFs/HTML, convert to
clean JSON with LLM-generated schema, publish as raw JSON in a public
GitHub repo (bypassing Google indexing entirely by letting AI crawlers
ingest it directly), then wrap the dataset as a free MCP server that
charges per-call via a usage-metered API key (RapidAPI free tier as
the billing layer).
2. Tech Stack: Python (requests/BeautifulSoup or Jina Reader prefix),
Google Colab (free compute), Gemini free-tier API (JSON structuring +
Schema.org generation), GitHub public repo (hosting),
modelcontextprotocol/servers template (MCP wrapper), RapidAPI free tier
(metered monetization).
3. Exact Steps:
a. `curl https://r.jina.ai/https://[county-site]/zoning-codes.pdf`
→ get clean markdown/text extraction free.
b. Prompt Gemini: "Convert this raw zoning text into a JSON array with
fields: parcel_id, zone_type, allowed_use, max_height_ft, source_url.
Return ONLY valid JSON, no prose."
c. Prompt again: "Generate a Schema.org JSON-LD block of type
GovernmentService for each entry above."
d. Commit `zoning_data.json` + `schema.jsonld` to a public GitHub repo
named descriptively (e.g., `county-zoning-open-data`).
e. Add a `README.md` with dense, information-gain-rich prose (unique
stats not published elsewhere) — this becomes the LLM citation
target.
f. Clone `modelcontextprotocol/servers/src/sqlite` template, point it
at a SQLite conversion of the JSON, deploy on a free Render/Fly.io
instance, list it on the MCP servers directory for Cursor/Cline
users, gate advanced queries behind a free RapidAPI key with a
pay-per-call tier above N free calls/month.
Example 2: Input: "Give me a free GEO tracking hack, no budget for LLMrefs Pro."
Output:
BUCKET: Free GEO Share-of-Voice Tracking
1. Tactical Play: Use crawler user-agent log analysis as a proxy for GEO
performance instead of paid SoV tools — if AI bots are hitting your
pages, you're in the training/retrieval pipeline.
2. Tech Stack: Vercel or Netlify (free hosting + free request logs),
a spreadsheet (Google Sheets), manual LLM query panel (free ChatGPT/
Perplexity/Gemini accounts).
3. Exact Steps:
a. In Vercel/Netlify dashboard, filter deployment logs for user-agents:
`GPTBot`, `OAI-SearchBot`, `PerplexityBot`, `ClaudeBot`, `Google-Extended`.
b. Log daily hit counts per page into a Sheet — rising hits = rising
GEO relevance.
c. Weekly, manually query 10-20 target prompts across ChatGPT/
Perplexity/Gemini; log whether your domain is cited verbatim,
linked, or absent — a manual, free FreeSOV equivalent.
d. Cross-reference: pages with bot hits but no citations need more
"Information Gain" injection (Step 3 of Flip phase).
- Always default to public-record / permissively licensed data — this is the highest-leverage, lowest-legal-risk ingestion target.
- Favor raw
.md/.jsonin public GitHub repos over full websites when the goal is pure LLM ingestion — it's faster to deploy and directly crawlable, no build step needed. - Every "Flip" step should explicitly prompt for Information Gain (facts/numbers not in the top existing sources) — this is what makes content citable by LLMs, not just SEO-optimized.
- Treat JSON-LD Schema.org markup as mandatory, not optional, in every publish step — it's free and materially increases structured citability.
- Stack monetization hooks (lead-gen form + MCP wrapper + affiliate links) rather than picking one — zero marginal cost to add more.
- Recommend rate-limited, polite scraping (respect robots.txt for scraping targets even while exploiting AI crawler behavior for distribution) to keep the loop sustainable and undetected.
- Don't fabricate named tools/posts/URLs that can't be verified — describe verifiable mechanisms (documented crawler user-agents, documented free-tier limits) instead of invented sources.
- Don't recommend scraping gated/paywalled/ToS-prohibited content as a "free" tactic — flag legal risk explicitly.
- Don't stop at theory — every bucket needs a runnable script snippet, prompt string, or config, not just a concept description.
- Don't conflate SEO tactics with GEO tactics — GEO cares about LLM citation/retrieval behavior, not SERP ranking; keep tracking metrics AI-crawler-specific.
- Don't recommend paid-tier tools as "free" — verify the specific tier (e.g., "Firecrawl free tier: 500 credits," not just "Firecrawl").