Architecting Cross Referenced Knowledge Bases
YAML--- name: architecting-cross-referenced-knowledge-bases description: Transforms disorganized collections of documents, notes, and research artifacts into interlinked, navigable knowledge systems using advanced hyperlinking strategies (bidirectional links, tagging taxonomies, semantic clustering, and index/hub pages). Use when a user has messy notebooks, sprawling document sets, research archives, or note collections that need to be restructured into a coherent, cross-referenced, browsable system. ---
Given a messy set of documents/notebooks, produce:
- A taxonomy (categories + tags) derived from actual content, not assumed structure.
- An index/hub document that links to every item, grouped by taxonomy.
- Cross-links between related items embedded inline (not just in the index).
- A naming/ID convention so links stay stable as content evolves.
Minimal output skeleton:
Markdown# Master Index
- [[Document Title 1]] — one-line description — related: [[Doc 3]], [[Doc 7]]
- [[Document Title 2]] — one-line description
...
- [[Untitled notebook (id: 4471)]]
Progress:
- Step 1: Inventory all items (title, id/url, rough content summary)
- Step 2: Cluster items by topic/theme (bottom-up, not forced into a preconceived scheme)
- Step 3: Name clusters clearly; flag ambiguous or duplicate items
- Step 4: Build hub pages per cluster with descriptive links
- Step 5: Add cross-links between related items across clusters (the "intricate hyperlinking" layer)
- Step 6: Build a top-level master index linking all hubs
- Step 7: Identify and flag orphans, duplicates, and "Untitled notebook" placeholders for cleanup
- Step 8: Verify every item is reachable from the master index in ≤2 clicks
Step 1 — Inventory: Extract every distinct title/item. Do not skip "Untitled notebook" entries — flag them explicitly as needing renaming.
Step 2 — Clustering: Group by actual subject overlap (e.g., "crochet pattern generation," "AI agent skill frameworks," "LLM research"). Expect items to belong to multiple clusters — this is where cross-linking earns its value.
Step 3 — Naming: Cluster names should be specific and scannable, not generic ("Crochet Pattern Generation & Chart Engines" not "Crafts").
Step 4 — Hub construction: Each hub is a page listing member items with one-line descriptions. Hubs are themselves linkable nodes.
Step 5 — Cross-linking: For each item, ask "what 2-4 other items does this relate to, and why?" Add inline related: links. This is the differentiator between a flat list and a true knowledge graph.
Step 6 — Master index: One page, linking to every hub, plus a search/lookup aid (alphabetical list or tag index).
Step 7 — Flag debt: Untitled items, duplicates (e.g., multiple "Wired USA" issues, multiple crochet-engine variants), and near-duplicate topics get a dedicated "Needs Triage" section rather than being silently dropped.
Step 8 — Verify reachability: Walk the graph from the master index; anything requiring 3+ hops gets an extra shortcut link.
Example 1: Input: 20 notebooks about crochet: "CrochetBench," "CrochetPARADE," "The Complete Crochet Compendium," "UGAFE Framework," "Multimodal Deterministic Crochet Pattern Compiler," "Go Gopher Amigurumi Pattern," etc. Output:
Markdownundefined
- [[CrochetBench]] — benchmark for generative crochet design — related: [[CrochetPARADE]], [[Multimodal Deterministic Crochet Pattern Compiler]]
- [[CrochetPARADE]] — pattern renderer/analyzer/debugger — related: [[CrochetBench]], [[UGAFE Framework]]
- [[UGAFE Framework]] — symbol parsing for crochet charts — related: [[CrochetPARADE]], [[The Complete Crochet Compendium]]
- [[The Complete Crochet Compendium]] — terms/techniques reference — related: [[UGAFE Framework]], [[Beginner's Guide to Granny Square]]
- [[Go Gopher Amigurumi Pattern]] — related: [[Multimodal Deterministic Crochet Pattern Compiler]] (uses same chart engine)
**Example 2:**
Input: Mixed bag — "Long-Term Memory Systems for AI Agents," "Agentic AI Workflows and LLMOps," "AI Agent Frameworks and Repository Maintenance Workflows," "Claude Code Engineering and Intelligence Skills."
Output: A single hub "AI Agent Architecture & Tooling" containing all four, with cross-links noting that memory systems feed into agentic workflows, which are implemented via the repository maintenance frameworks, which are exercised by Claude Code skills — i.e., link chains follow the actual dependency/usage relationship, not alphabetical order.
- Derive categories from content, never impose a template before reading titles/summaries.
- Prefer many small, precise hubs over one giant miscellany page.
- Every cross-link should state why (short parenthetical), not just exist.
- Use stable identifiers (slugs or IDs) for links so renaming a title doesn't break the graph.
- Explicitly surface duplicates/near-duplicates instead of silently merging or ignoring them.
- Keep the master index skimmable — descriptions max ~12 words.
- Don't create a flat alphabetical list and call it "organized" — that's not hyperlinking, it's sorting.
- Don't leave "Untitled notebook" entries unflagged; they will silently become orphans.
- Don't force every item into exactly one category — real content overlaps.
- Don't add links without a stated relationship — meaningless links add noise, not navigability.
- Don't let hub pages balloon past ~15-20 items; split into sub-hubs instead.