AI Skill Report Card

Summarizing Documents

A-86·Sep 28, 2026·Source: Web
14 / 15

Given a document file, follow this pipeline:

1. Check file type → must be PDF, DOCX, or PPTX (else reject)
2. Convert to Markdown using doc2md
3. Inspect the Markdown for gaps (missing pages, unreadable text, broken tables)
4. Summarize the Markdown content only — never add outside facts
5. Save: <filename>.md (source) + <filename>-summary.md (summary)
6. Report: filename, conversion status, any flagged issues

Output format for the summary file:

Markdown
# Summary: <original filename> **Conversion status:** [clean | partial — see flags below]
Recommendation▾
Clarify what 'doc2md' is—specify whether it's a bundled script, CLI tool, or needs installation, since Claude may not have this tool available
  • ...

<Section/Heading from source>

  • ...
  • [e.g., "Table on page 4 did not convert; verify against original."]
14 / 15

Progress:

  • Step 1: Validate file type (PDF, DOCX, PPTX only)
  • Step 2: Convert file to Markdown with doc2md
  • Step 3: Inspect converted Markdown for missing/garbled content
  • Step 4: Summarize using only extracted content
  • Step 5: Save source Markdown + summary side-by-side with original
  • Step 6: Report filename, conversion status, and flags

Step 1 — Validate: If the file isn't PDF/DOCX/PPTX, stop and report "unsupported format," don't attempt conversion.

Step 2 — Convert: Run doc2md on the file. Capture any tool errors verbatim rather than suppressing them.

Step 3 — Inspect before summarizing: Read the raw Markdown output first. Note:

  • Blank or near-empty pages/slides (likely scanned images without OCR)
  • Tables rendered as broken/misaligned text
  • Truncated sections This step is separate from summarization — extraction problems must be caught before they get baked into a "confident" summary.

Step 4 — Summarize:

  • Preserve: main claims, evidence, names, dates, figures, units, conclusions.
  • Use headings/bullets mirroring the source structure.
  • For PPTX, keep slide titles as headings; add one overall-argument summary plus per-slide points.
  • For long documents, summarize section-by-section first, then compose an overall summary from those.
  • Never invent details. If extracted text is ambiguous, mark it as uncertain rather than presenting it as fact.

Step 5 — Save outputs: Place both the raw Markdown and the summary next to the original file, using consistent naming.

Step 6 — Report: Always state original filename, conversion status (clean/partial/failed), and any flagged gaps.

Recommendation▾
Add an edge case for password-protected or corrupted files that can't even be opened
19 / 20

Example 1: Input: quarterly-report.pdf (text-based PDF, one table on page 4 fails to convert) Output:

Markdown
# Summary: quarterly-report.pdf **Conversion status:** partial
Recommendation▾
Include guidance on handling very large documents that might exceed context limits during the section-by-section summarization step
  • Revenue grew 12% YoY, driven by APAC expansion.
  • Operating costs rose 8%, mainly from hiring.

Financial Overview

  • Q3 revenue: $4.2M (source states figure directly)
  • Net margin: 18%
  • Table on page 4 (regional breakdown) did not convert cleanly — verify against original PDF.

**Example 2:**
Input: `product-pitch.pptx` (12 slides, brief bullet text per slide)
Output:
```markdown
# Summary: product-pitch.pptx

**Conversion status:** clean

The deck argues that mid-market SaaS teams need a lighter-weight analytics tool than current enterprise options, positioning the product as the "80% solution at 20% of the cost."

Slide 3: The Problem

  • Enterprise analytics tools are over-built for mid-market teams.

Slide 7: Pricing

  • Three tiers: $29/$99/$299 per month.
  • None.

**Example 3:**
Input: `scanned-contract.pdf` (image-only pages, no OCR layer)
Output:
```markdown
# Summary: scanned-contract.pdf

**Conversion status:** failed
  • Document appears to be scanned images with no extractable text layer. OCR is required before summarization can proceed. No summary generated to avoid fabricating content.
  • Treat extraction and summarization as separate, sequential steps — always inspect converted Markdown before summarizing it.
  • Keep source Markdown alongside the summary so claims can be traced back.
  • For tables, preserve labels and units exactly; don't recompute or reformat numbers.
  • When a document is long, summarize per-section, then synthesize — don't try to summarize a huge blob in one pass.
  • Always include filename and conversion status in every output, even successful ones.
  • Don't summarize past a conversion failure — if extraction is empty or garbled, report the failure instead of guessing content.
  • Don't silently drop tables or figures that failed to convert — flag them explicitly.
  • Don't treat OCR-free scanned PDFs as equivalent to text PDFs; check for empty extraction first.
  • Don't add context, interpretation, or outside knowledge not present in the source text.
  • Don't collapse slide-specific nuance in PPTX summaries into only a general overview — include both levels.
0
Grade A-AI Skill Framework
Scorecard
Criteria Breakdown
Quick Start
14/15
Workflow
14/15
Examples
19/20
Completeness
17/20
Format
14/15
Conciseness
14/15