Inventorying Chat Archives
Given a raw scrape of chat/notebook titles (often duplicated, jammed together, or interspersed with UI chrome like "Open sidebar" or "New thread"), produce a clean manifest table:
Markdown# Chat Archive Manifest | # | Title | Category | Notes | |---|-------|----------|-------| | 1 | Global Frameworks for AI Governance and Educational Integrity | AI Governance & Policy | | | 2 | Vintage Crochet Pattern Archive: Afghans, Appliques, and Baby Essentials | Crochet/Craft Patterns | | | 3 | Data-Exhaust & Market Opportunity Grounding Repository | Business/Market Research | |
No preamble — start extracting immediately.
Progress:
- Step 1: Strip UI noise (nav labels like "Open sidebar", "New thread", "Manage AI Mode shared public links", "Search threads", "Close sidebar", "Personalization", "Settings")
- Step 2: De-duplicate — titles are typically repeated twice in a row (link text + accessible text); collapse each pair into one entry
- Step 3: Normalize "Untitled notebook" entries — count them but list as a single grouped row (e.g., "Untitled notebook ×12") rather than 12 separate rows, unless the user wants every instance tracked individually
- Step 4: Assign each unique title a sequential ID number
- Step 5: Infer a category per title based on keywords (see Category Heuristics below)
- Step 6: Group manifest by category OR keep flat numbered list with a category column — default to flat list with category column for easy sorting/searching
- Step 7: Add summary statistics block at top (total count, count per category, count of untitled/generic entries)
- Step 8: Output final manifest as a markdown table (or CSV if requested)
Infer categories from title keywords — don't ask the user, just apply best judgment:
- Crochet/Craft Patterns: crochet, amigurumi, knitting, beading, tatting, textile, weaving, granny square, yarn
- AI/LLM Engineering: LLM, agent, Claude, Gemini, MCP, prompt, skill, LLMOps, hallucination, RAG, embeddings
- Software/Dev Tools: GitHub, repository, COBOL, compiler, AST, monorepo, API, vanilla JS
- Research/Academic: arXiv, research papers, physics, mathematics, benchmark, evaluation
- Data/Database: dataset, Parquet, vector database, DB2, CSV/YAML indices
- Business/Legal/Policy: governance, ethics, utility subsidies, patents, forensic
- Personal/Misc: sleep hygiene, personas, SMS logs, deleted items
- Untitled/Unclassified: generic "Untitled notebook", "export", "drive", single ambiguous words
If a title fits multiple categories, pick the most specific one and note ambiguity in the Notes column rather than agonizing over it.
Example 1:
Input: Untitled notebookUntitled notebookThe Go Gopher Amigurumi Pattern and GitHub Development SolutionsThe Go Gopher Amigurumi Pattern and GitHub Development Solutions
Output:
| # | Title | Category |
|---|---|---|
| — | Untitled notebook ×1 (counted in summary) | Unclassified |
| 1 | The Go Gopher Amigurumi Pattern and GitHub Development Solutions | Crochet/Craft Patterns + Dev Tools |
Example 2: Input: A 150-title dump mixing crochet notebooks, AI agent frameworks, and research indices.
Output: A manifest with a summary header:
Total unique threads: 147
- Crochet/Craft Patterns: 34
- AI/LLM Engineering: 41
- Research/Academic: 22
- Software/Dev Tools: 18
- Untitled/Unclassified: 15
- Other: 17
...followed by the full numbered table.
- Always de-duplicate the doubled title pattern first — this is the most common artifact from scraped web UIs.
- Collapse repetitive "Untitled notebook" noise into a count rather than cluttering the manifest.
- Preserve exact original title text (don't paraphrase) — this is an inventory, not a summary.
- Output a summary stats block before the full table so the user gets immediate value even for huge lists.
- If the list exceeds ~200 entries, offer to split output into multiple files/sections by category.
- Default to a flat table with a Category column (sortable) rather than nested category sections, unless user explicitly wants grouped output.
- Don't ask the user to clarify categories before producing output — infer and label, note uncertainty in Notes column instead.
- Don't drop entries silently when unsure of category — always bucket into "Other/Unclassified" rather than omitting.
- Don't merge distinct titles that are merely similar (e.g., "GitDiagram Public Repository Catalog" vs "GitDiagram Browser Repository Links" are separate entries).
- Don't retain UI chrome text as if it were a real thread title.
- Don't renumber or reorder arbitrarily between revisions — keep IDs stable if the manifest is iterated on.