Extracting Structured Data From Unstructured Text
Given raw pasted text and a target data type, immediately identify the extraction target, scan for pattern signals, and return ONLY the structured result — no restated input, no commentary unless asked.
Input: [3,000-word HathiTrust OCR dump]
Target: granular index headings
Output:
- Agriculture, Dept. of — see also Farm Bureau
- Extension Service, 112-118
- Soil Conservation, 119-124
- Agriculture, State Board of, 45-50
- Identify the extraction target — the noun phrase in the request (headings, call numbers, coordinates, ratios, addresses, milestones, etc.). This defines the schema of the output.
- Identify the source noise type — OCR garble, chat log timestamps, marketing fluff, metadata key-value soup, stitch abbreviations, CSS blocks, frontmatter, URLs. This defines what to discard.
- Scan for pattern anchors — delimiters, repeated punctuation, numeric ranges, known abbreviation sets, indentation/hierarchy cues, known field labels (e.g., "Call No:", "Shift:", "Tie-up:").
- Extract and normalize — pull only the matching spans; normalize formatting (consistent casing, units, delimiters) without altering the underlying values.
- Validate against domain constraints — if the domain has known invariants (binary tie-up matrix must be 0/1, stitch counts must sum correctly, CSS values must be valid units, addresses must contain city/state/zip), check and flag violations rather than silently "fixing" them.
- Output only the structured result in the simplest faithful format (list, table, key-value, code block) — match output shape to what a downstream tool/human would consume (e.g., clipboard-paste-ready for CSS/texture params, plain sentences for stitch counts).
Progress checklist for multi-part or long inputs:
- [ ] Target schema identified
- [ ] Noise type identified
- [ ] Pattern anchors located
- [ ] Extraction pass complete
- [ ] Domain validation pass complete
- [ ] Output formatted for intended downstream use
Example 1 — Weaving tie-up verification
Input: Tie-up: 1000 0100 0010 0001 (4 shafts x 4 treadles claimed)
Output:
Shaft 1 → Treadle 1
Shaft 2 → Treadle 2
Shaft 3 → Treadle 3
Shaft 4 → Treadle 4
Matrix valid: 4x4, one tie per treadle, no conflicts.
Example 2 — Stitch abbreviation to plain English
Input: k2, p1, k2tog, yo, k1, ssk, p2
Output: Knit 2, purl 1, knit 2 together, yarn over, knit 1, slip-slip-knit, purl 2. Row stitch count change: -0 (one decrease offset by one increase).
Example 3 — Redacting cloud-dependent marketing phrasing Input: "Access your patterns anywhere with our cloud-synced app and real-time online dashboard!" Output: "Access your patterns anytime with the included offline app." (Flagged/removed: "cloud-synced", "online dashboard")
Example 4 — Business address isolation from messy confirmation text Input: "Confirmed! Job #4471 — Mike's Hardware, checked in at 3:42pm, addr looks like 118 Elm St Unit B, Springfield IL 62701, allowance $12 base fee $8.50" Output:
Business: Mike's Hardware
Address: 118 Elm St, Unit B, Springfield, IL 62701
Allowance: $12.00
Base fee: $8.50
Net payout: $20.50
- Preserve original values exactly (numbers, codes, units) — never round, reformat units, or "correct" spelling inside extracted data unless explicitly asked to normalize.
- When hierarchy exists (indexes, CSS nesting, course grids), preserve nesting/order in output.
- When a domain has a known invariant (tie-up matrix squareness, stitch count parity, sum of hours in a shift), state whether it passed or failed — don't just extract silently.
- Default to the leanest output format usable by a machine or copy-paste workflow (comma-separated, key:value, code block) unless prose is explicitly requested (e.g., "plain-English counts").
- If the source text contains multiple candidate matches for the same field (e.g., two dollar amounts), surface all candidates with brief context rather than guessing.
- If extraction target is ambiguous, extract the most literal/narrow interpretation first, then note broader candidates.
- Don't return the entire cleaned-up input when only a subset was requested — extract, don't reformat everything.
- Don't infer missing structured fields (e.g., a missing zip code) — flag as missing, don't fabricate.
- Don't silently drop validation failures (e.g., an invalid tie-up matrix, a stitch count that doesn't add up) — surface them.
- Don't over-explain the extraction process unless asked; default output is the data itself.
- Don't apply generic NLP cleanup (stripping all punctuation) if it destroys domain-meaningful syntax (e.g., WIF section brackets, CSS units, stitch abbreviation commas).