Parsing and Synthesizing Technical Documents
Given raw source material (papers, docs, transcripts, multiple articles), produce a ~2000-word synthesis:
- Extract: claims, definitions, methods, data points, contradictions
- Cluster: group extracted items into 4-6 thematic sections
- Synthesize: write connected prose per section, not bullet dumps
- Verify: every claim traceable to source; no invented facts
Output structure:
# Title
Progress:
- Step 1: Inventory source material — list every distinct document/section and its core topic
- Step 2: Extract atomic units — pull out discrete facts, definitions, numbers, methods, quotes with locators (page/section/paragraph)
- Step 3: Deduplicate and reconcile — merge repeated points; flag contradictions explicitly rather than silently picking one
- Step 4: Build thematic clusters — group atomic units by concept, not by source document
- Step 5: Rank by density/importance — cut low-value tangents; keep only what changes understanding or action
- Step 6: Draft synthesis prose — write connected paragraphs that explain relationships between facts (causal, comparative, sequential)
- Step 7: Compress to target length — enforce word budget per section; cut filler, keep specifics (numbers, names, mechanisms)
- Step 8: Fact-check against source — reread draft against original; strike anything unverifiable
- Step 9: Add attribution — note which source(s) each major claim traces to
Example 1: Input: A 40-page technical report on brown dwarf radius inflation with 15 cited studies, mixed observational data and theoretical models. Output: A 2000-word document with sections: Overview of the inflation puzzle; Observational evidence (specific radii, masses, systems named); Competing theoretical mechanisms (tidal heating, magnetic activity, metallicity) with the actual physics distinguishing them; Points of disagreement between studies; Implications for future surveys. Each theoretical mechanism gets 2-3 sentences explaining why it produces inflation, not just that it's proposed.
Example 2: Input: Three overlapping API documentation pages for a vector database (architecture, performance benchmarks, implementation guide). Output: One document merging them: Architecture (index types, storage layout); Performance characteristics (actual benchmark numbers, tradeoffs by index type); Implementation guidance (concrete config choices mapped to use cases); a short table contrasting index types by latency/memory/recall. Redundant boilerplate across the three pages is collapsed into single statements.
- Density over coverage. Cut anything a domain expert would consider filler (generic intros, repeated definitions, marketing language).
- Preserve specifics. Numbers, names, formulas, version numbers, dates — these are what make a document "useful." Never round or vague them out.
- Show relationships, not lists. Prefer "X causes Y because Z" over "X. Y. Z." Bullet lists are for enumerable items only (steps, parameters), not for explanatory content.
- Flag uncertainty explicitly. If sources disagree or a claim is weakly supported, say so in one clause rather than smoothing it over.
- Use a consistent unit of extraction. Track claims as (assertion, source, confidence) triples internally even if not shown in output.
- Match structure to content type. Empirical material → evidence/mechanism/implication structure. Procedural material → sequential steps. Comparative material → tables.
- Padding to hit word count with restated intros/conclusions — cut a weak section instead of padding a strong one.
- Summarizing source-by-source instead of theme-by-theme — produces a disjointed patchwork, not a synthesis.
- Dropping all numbers/specifics in favor of "high-level" prose — this destroys the density the task requires.
- Inventing connective claims not actually in the source ("this suggests...") when the source never draws that inference — mark inferences as your own, distinct from source claims.
- Treating every source sentence as equally important — always rank and cut.