AI Skill Report Card

Extracting Structured Data From Unstructured Text

A-84·Sep 29, 2026·Source: Web
14 / 15

Given raw pasted text and a target data type, immediately identify the extraction target, scan for pattern signals, and return ONLY the structured result — no restated input, no commentary unless asked.

Input: [3,000-word HathiTrust OCR dump]
Target: granular index headings
Output:
- Agriculture, Dept. of — see also Farm Bureau
  - Extension Service, 112-118
  - Soil Conservation, 119-124
- Agriculture, State Board of, 45-50
Recommendation▾
Add a 'bad output' counterexample (e.g., an over-verbose or hallucinated extraction) to more explicitly contrast good vs. bad outcomes
14 / 15
  1. Identify the extraction target — the noun phrase in the request (headings, call numbers, coordinates, ratios, addresses, milestones, etc.). This defines the schema of the output.
  2. Identify the source noise type — OCR garble, chat log timestamps, marketing fluff, metadata key-value soup, stitch abbreviations, CSS blocks, frontmatter, URLs. This defines what to discard.
  3. Scan for pattern anchors — delimiters, repeated punctuation, numeric ranges, known abbreviation sets, indentation/hierarchy cues, known field labels (e.g., "Call No:", "Shift:", "Tie-up:").
  4. Extract and normalize — pull only the matching spans; normalize formatting (consistent casing, units, delimiters) without altering the underlying values.
  5. Validate against domain constraints — if the domain has known invariants (binary tie-up matrix must be 0/1, stitch counts must sum correctly, CSS values must be valid units, addresses must contain city/state/zip), check and flag violations rather than silently "fixing" them.
  6. Output only the structured result in the simplest faithful format (list, table, key-value, code block) — match output shape to what a downstream tool/human would consume (e.g., clipboard-paste-ready for CSS/texture params, plain sentences for stitch counts).

Progress checklist for multi-part or long inputs:

- [ ] Target schema identified
- [ ] Noise type identified
- [ ] Pattern anchors located
- [ ] Extraction pass complete
- [ ] Domain validation pass complete
- [ ] Output formatted for intended downstream use
Recommendation▾
The domain examples (weaving, knitting) are oddly specific/niche relative to the generic description — consider broader examples like log parsing or JSON extraction to match the stated scope
17 / 20

Example 1 — Weaving tie-up verification Input: Tie-up: 1000 0100 0010 0001 (4 shafts x 4 treadles claimed) Output:

Shaft 1 → Treadle 1
Shaft 2 → Treadle 2
Shaft 3 → Treadle 3
Shaft 4 → Treadle 4
Matrix valid: 4x4, one tie per treadle, no conflicts.

Example 2 — Stitch abbreviation to plain English Input: k2, p1, k2tog, yo, k1, ssk, p2 Output: Knit 2, purl 1, knit 2 together, yarn over, knit 1, slip-slip-knit, purl 2. Row stitch count change: -0 (one decrease offset by one increase).

Example 3 — Redacting cloud-dependent marketing phrasing Input: "Access your patterns anywhere with our cloud-synced app and real-time online dashboard!" Output: "Access your patterns anytime with the included offline app." (Flagged/removed: "cloud-synced", "online dashboard")

Example 4 — Business address isolation from messy confirmation text Input: "Confirmed! Job #4471 — Mike's Hardware, checked in at 3:42pm, addr looks like 118 Elm St Unit B, Springfield IL 62701, allowance $12 base fee $8.50" Output:

Business: Mike's Hardware
Address: 118 Elm St, Unit B, Springfield, IL 62701
Allowance: $12.00
Base fee: $8.50
Net payout: $20.50
Recommendation▾
Clarify output format decision rules more concretely (e.g., a small decision table for when to use table vs. key-value vs. list) rather than leaving it to judgment
  • Preserve original values exactly (numbers, codes, units) — never round, reformat units, or "correct" spelling inside extracted data unless explicitly asked to normalize.
  • When hierarchy exists (indexes, CSS nesting, course grids), preserve nesting/order in output.
  • When a domain has a known invariant (tie-up matrix squareness, stitch count parity, sum of hours in a shift), state whether it passed or failed — don't just extract silently.
  • Default to the leanest output format usable by a machine or copy-paste workflow (comma-separated, key:value, code block) unless prose is explicitly requested (e.g., "plain-English counts").
  • If the source text contains multiple candidate matches for the same field (e.g., two dollar amounts), surface all candidates with brief context rather than guessing.
  • If extraction target is ambiguous, extract the most literal/narrow interpretation first, then note broader candidates.
  • Don't return the entire cleaned-up input when only a subset was requested — extract, don't reformat everything.
  • Don't infer missing structured fields (e.g., a missing zip code) — flag as missing, don't fabricate.
  • Don't silently drop validation failures (e.g., an invalid tie-up matrix, a stitch count that doesn't add up) — surface them.
  • Don't over-explain the extraction process unless asked; default output is the data itself.
  • Don't apply generic NLP cleanup (stripping all punctuation) if it destroys domain-meaningful syntax (e.g., WIF section brackets, CSS units, stitch abbreviation commas).
0
Grade A-AI Skill Framework
Scorecard
Criteria Breakdown
Quick Start
14/15
Workflow
14/15
Examples
17/20
Completeness
17/20
Format
15/15
Conciseness
13/15