Architecting Atomic AI Task Systems
Given any candidate task (a prompt, a workflow step, or a described process), run it through the 5-Point Atomicity Acid Test before building anything:
- Single I/O — exactly one input type, one output type
- Zero branching — no internal if/else/routing logic
- Independent testability — one (input, output) pair fully validates it
- Conjunction-free scope — one sentence, no "and"
- Decoupled orchestration — no fetches, writes, emails, or routing inside the task
If it fails any point, split along the failing axis and re-test each fragment. Only atomic tasks get built as prompts; everything else (fetching, routing, retries, persistence) belongs to the orchestrator layer.
Progress:
- Step 1: Define the task's one-sentence functional scope (reject if it contains "and"/"or")
- Step 2: Run the 5-Point Atomicity Acid Test; split if any point fails
- Step 3: Classify every operation as Cognitive Layer (LLM) or Orchestrator Layer (system action) — never mix
- Step 4: Write the full task spec (input schema, output schema, edge cases, failure flags, test vector)
- Step 5: Assign standardized error enums (
REJECT_*,FLAG_*,ERR_*) for every anomaly path - Step 6: If extracting from a raw source (SOP, checklist, process doc, prompt library), apply the matching 3-part extraction template (IDENTIFY → CONVERT → OUTPUT)
- Step 7: Compose atomic tasks into a pipeline via orchestrator-owned chaining, fan-out/fan-in, and gates — not via task-internal logic
- Step 8: Validate the full chain against a static test vector before deployment
Every atomic task entry must contain:
Task label & canonical name (gerund or verb-noun form, e.g., isolate-row-stitch-counts)
Functional scope: one sentence, no "and"
Input schema: MIME type, size/shape bounds, concrete example
Output schema: MIME type, shape, concrete example
Edge cases: explicit empty/malformed/boundary behaviors
Failure flags: which REJECT_*/FLAG_*/ERR_* apply and what the task does in each case
Test vector: one static (input, expected_output) pair
Never let a task perform: fetching URLs, reading/writing files, querying databases, sending notifications, retrying, fan-out/fan-in, or conditional routing based on its own output. These are always orchestrator responsibilities. A task that says "fetch X then parse Y" is two things wearing a trenchcoat — split it.
| System Action | Owner | Never inside |
|---|---|---|
| fetch-url, send-email, write-to-db, trigger-webhook, route-if-flag, retry-on-failure, fan-out/fan-in | Orchestrator | Any atomic task |
Use exactly three prefixes, no exceptions:
REJECT_*— task aborts, produces no output (e.g.,REJECT_INSUFFICIENT_INPUT,REJECT_MALFORMED_TUPLE,REJECT_INVARIANT_ERROR)FLAG_*— task produces degraded/partial output with an annotation (e.g.,FLAG_NO_TARGET_DETECTED,FLAG_AMBIGUOUS_ACTOR,FLAG_OUT_OF_RANGE_VALUE)ERR_*— arithmetic or invariant computation failure (e.g.,ERR_ARITHMETIC_INVARIANT_MISMATCH,ERR_UNVERIFIED_HHI)
Rule: tasks never retry (orchestrator's job) and never silently swallow errors — every anomaly surfaces as a flag or a null-valued field.
Every extraction prompt has exactly three parts: IDENTIFY (numbered, schema-specific structural elements — never generic), CONVERT (explicit target schema + mapping rule), OUTPUT (exact field names, enums, ${variable} / ${variable:default} placeholder syntax).
Pick the matching domain schema:
- SOP → Sequential Execution Chain: TRIGGER → ROLE → STEP_SEQUENCE → VERIFICATION → OUTPUT, with output_artifact binding step-to-step
- Checklist → State Machine: states + transitions + guards, e.g.
(from_state, event, guard, to_state) - Process Doc → Gated Routing Graph: nodes/edges/gates/exception handlers
- Prompt Library → YAML Config: contexts + roles + tasks + formats + constraints, composed via
${required}/${optional:default}
Example 1:
Input: "Analyze the transcript, summarize the complaint, and route it to the right team."
Output: Fails points 2 and 4 (branching + conjunction). Split into three atomic tasks: extract-complaint-summary (text→text), classify-complaint-category (text→enum), with routing itself left entirely to the orchestrator based on the enum output.
Example 2:
Input: A craft pattern text block: "R1: (6). R2: inc each st (12). R3: sc, inc x6 (18)."
Output: isolate-row-stitch-counts task spec → {"stitch_counts":[6,12,18]}, with REJECT_NON_INTEGER_TOKEN firing per malformed token and empty input yielding {"stitch_counts":[]}.
- Name tasks as verb-noun or gerund phrases describing one transformation only (
isolate-row-stitch-counts, notprocess-pattern) - Always write the test vector before writing the task's full spec — if you can't produce one static example pair, the task isn't atomic yet
- Prefer many small tasks chained by the orchestrator over one flexible task with internal logic
- Every optional parameter uses
${name:default}; every required parameter uses${name}with no default - When in doubt about ambiguous input, emit a
FLAG_*with a best-effort partial output rather than a hardREJECT_*
- Hidden branching: "if it's a complaint, extract X, otherwise extract Y" is two tasks disguised as one — split immediately
- Conjunction creep: "extract and rank" or "fetch and summarize" always fails the acid test — the "and" is the tell
- Orchestration leakage: a task that calls an API or writes a file is no longer atomic and can't be unit-tested with a static pair
- Silent failure: returning an empty result with no flag hides real errors from the orchestrator — always emit the specific
FLAG_*/REJECT_* - Retry logic inside a task: retries are stateful and belong exclusively to the orchestrator layer