AI Skill Report Card

Extracting Atomic Tasks

A89·Sep 29, 2026·Source: Web

Decomposes messy real-world workflows (sales playbooks, SOPs, checklists, process docs) into atomic, independently testable task units with formal specifications, ready to be assembled into pipelines by an orchestrator layer.

14 / 15

Given a corpus of workflow documents, produce entries like this for each atomic task:

YAML
task_id: extract-lead-company-name description: > Identify the company name mentioned in a raw outreach lead note. input: type: text/plain schema: { max_chars: 2000 } output: type: application/json schema: { company_name: string } errors: - case: empty_input flag: REJECT_EMPTY_INPUT - case: no_company_found flag: FLAG_NO_ENTITY - case: multiple_candidates flag: FLAG_AMBIGUOUS_ENTITY bricks: context: "Raw CRM lead note, unstructured" role: "Entity extraction specialist" task: "Identify → Convert → Output company name" format: "JSON: {company_name: string}" constraint: "No inference beyond explicit text; no external lookups" test_vectors: - input: "Met Jane from Acme Corp at conference." expected_output: { company_name: "Acme Corp" } - input: "" expected_output: { error: "REJECT_EMPTY_INPUT" }

Never include system actions (API calls, DB writes, notifications) inside a task spec — those belong to the orchestrator, not the atomic task.

Recommendation▾
Add a second full example showing a task that passes the acid test on first try (not just splitting examples), to balance the two given.
15 / 15

Work chunk-by-chunk. Never try to atomize an entire corpus at once — partition first, then decompose, then formalize, then bind.

Progress:
- [ ] Phase 1: Partition source corpus into chunks by domain
- [ ] Phase 2: Extract atomic tasks per chunk
- [ ] Phase 3: Author formal specs (schemas + error handling)
- [ ] Phase 4: Bind to reusable bricks, generate test vectors, compile registry

Phase 1 — Partition the corpus Group source material by domain/theme (e.g., Sales & Outreach, Finance & Legal, Careers & Productivity, SOP Repositories, Checklists). Estimate candidate task count per chunk to scope effort. Process one chunk fully before moving to the next.

Phase 2 — Extract atomic tasks

  1. System vs. Cognitive Split: Remove anything that touches the outside world (API calls, scraping, file I/O, sending messages). Only pure transformation/reasoning steps remain as candidate tasks.
  2. Primitive Decomposition: Force every remaining step through Identify → Convert → Output. If a step doesn't cleanly fit this anatomy, it's not atomic yet.
  3. Atomicity Acid Test: A candidate task passes only if it satisfies all 5:
    • Single input, single output
    • Zero branching/conditional logic inside the task
    • Independently testable in isolation
    • Verb-first name with no conjunctions ("and", "then", "or")
    • Fully decoupled from orchestration/sequencing concerns
  4. Iterative Splitting: Any task failing the acid test (e.g., analyze-and-summarize, validate-then-format) gets split into separate single-responsibility tasks. Re-run the acid test on each fragment.

Phase 3 — Formal specification

  1. Schema Binding: Nail down exact MIME type, character/size boundaries, and JSON/Markdown schema for both input and output. No free-form "text" without bounds.
  2. Boundary & Error Handling: Enumerate ≥3 edge cases per task (empty input, malformed tokens, missing required fields, ambiguous/multiple matches). Map each to a canonical flag using the naming convention:
    • REJECT_* — input cannot be processed at all
    • FLAG_* — input processed but with a caveat/ambiguity
    • ERR_* — internal/unexpected failure during transformation

Phase 4 — Bind & register

  1. CRTF / C-P-O Binding: Attach reusable CONTEXT / ROLE / TASK / FORMAT / CONSTRAINT bricks so the task spec is prompt-ready and composable with other tasks.
  2. Test Vector Generation: Write static (mock_input, expected_output) pairs covering the happy path and every documented error case.
  3. Registry Compilation: Append the validated spec to the Master Task Registry (YAML or JSONL), keyed by task_id.
Recommendation▾
Include a brief example of a bad/failing spec output (e.g., vague schema) alongside its corrected version to reinforce pitfalls concretely.
16 / 20

Example 1: Input: SOP step — "Review the contract, flag any missing indemnification clauses, and summarize the risk in plain English." Output: Rejected as atomic. Split into three tasks: detect-missing-clause (input: contract text + clause type; output: boolean + location), classify-clause-risk (input: clause text; output: risk level enum), summarize-risk-plain-english (input: risk level + clause context; output: 1-2 sentence string).

Example 2: Input: "Isolate GLOSOLAN soil procedure step: 'Record sample pH and reject readings outside 0–14.'" Output:

YAML
task_id: validate-ph-reading input: { type: application/json, schema: { ph_value: number } } output: { type: application/json, schema: { valid: boolean } } errors: - case: out_of_range flag: REJECT_OUT_OF_RANGE - case: non_numeric flag: ERR_MALFORMED_INPUT
Recommendation▾
Consider adding guidance on how to handle ambiguous domain boundaries during Phase 1 partitioning, since real corpora often mix domains.
  • Default to splitting when in doubt — an over-atomized task is easy to compose; an under-atomized one blocks reuse and testing.
  • Name tasks verb-noun (e.g., extract-entity, classify-sentiment), never verb-noun-and-verb-noun.
  • Keep orchestration concerns (retries, sequencing, branching between tasks) entirely out of task specs — that's the orchestrator's job.
  • Always produce at least one error-path test vector per task, not just the happy path.
  • Process one chunk to full completion (Phases 2–4) before starting the next chunk's extraction — don't batch Phase 1 across all chunks then stall.
  • Embedding system actions: A task that "looks up the company in Salesforce" is not atomic — it's orchestration + a separate format-lookup-result task.
  • Vague schemas: "Output: some text" is not a schema. Every field needs a type and boundary.
  • Skipping the acid test: Tasks that "mostly" fit still need all 5 criteria checked explicitly — silent branching logic is the most common failure mode.
  • Under-specifying errors: Only documenting the empty-input case misses malformed and ambiguous-input cases, which are the ones that break pipelines in production.
  • Conjunction names: If the task name needs "and"/"then" to describe it, it isn't done splitting.
0
Grade AAI Skill Framework
Scorecard
Criteria Breakdown
Quick Start
14/15
Workflow
15/15
Examples
16/20
Completeness
16/20
Format
14/15
Conciseness
14/15