AI Skill Report Card
Decomposing Tasks for Text Models
YAML--- name: decomposing-tasks-for-text-models description: Breaks down complex workflows into atomic, well-scoped tasks suitable for training or prompting focused text models. Use when a monolithic prompt, document set, or process needs to be decomposed into discrete, testable, single-responsibility units for model training, fine-tuning, or agent skill design. ---
Building a Text Model for Atomic Tasks
Quick Start13 / 15
Given a messy, multi-purpose document, prompt, or workflow, decompose it into atomic tasks:
- Identify the end-to-end goal.
- List every distinct sub-decision or transformation required to get there.
- For each sub-decision, write it as a single input → single output unit with no branching logic.
- Name each atomic task with a verb-first label (e.g.,
extract-entities,classify-intent,normalize-date).
Example decomposition:
Messy task: "Summarize this document and also tag it and also check for compliance issues."
Atomic tasks:
1. summarize-document (input: raw text -> output: summary)
2. tag-document (input: raw text -> output: tag list)
3. flag-compliance-issues (input: raw text -> output: issue list or none)
Recommendation▾
Add a third example showing a failed/bad decomposition (e.g., over-fragmented or still-branching task) alongside its corrected version to make the good vs. bad contrast more explicit.
Workflow14 / 15
Progress:
- Step 1: Define the overall objective and success criteria
- Step 2: Inventory all implicit sub-tasks embedded in the current process
- Step 3: Test each candidate task for atomicity (single input type, single output type, no conditional branching)
- Step 4: Split any task that fails the atomicity test
- Step 5: Define strict input/output schemas per task
- Step 6: Order tasks into a pipeline (sequential, parallel, or conditional routing)
- Step 7: Validate each atomic task independently with sample inputs
- Step 8: Assemble pipeline and validate end-to-end output against original objective
Atomicity test — a task is atomic if:
- It has exactly one clear input type and one clear output type.
- It requires no internal "if/else" business logic (routing belongs in the pipeline, not the task).
- It can be validated with a simple input/output pair without needing the rest of the pipeline.
- A human could describe its job in one sentence without using "and."
If a task fails any of these, split it further.
Recommendation▾
Include a concrete input/output schema example (e.g., JSON) for at least one atomic task to demonstrate the 'strict minimal output schema' guidance in practice.
Examples15 / 20
Example 1: Input: "Model that reads customer emails, decides sentiment, drafts a reply, and files it in the CRM." Output:
1. classify-sentiment (email text -> sentiment label)
2. extract-key-request (email text -> structured request object)
3. draft-reply (request object + sentiment -> reply text)
4. select-crm-category (request object -> category label)
(Filing itself is a system action, not a text-model task — excluded from the model's scope.)
Example 2: Input: A 40-page messy internal wiki mixing onboarding steps, troubleshooting FAQs, and policy text. Output:
1. classify-document-section (chunk -> {onboarding | faq | policy})
2. extract-step-sequence (onboarding chunk -> ordered step list)
3. extract-qa-pair (faq chunk -> question, answer)
4. extract-policy-rule (policy chunk -> rule statement, applicability)
Each becomes an independently testable atomic task/skill.
Recommendation▾
Slightly tighten the Workflow checklist by merging overlapping steps (e.g., 3 and 4 are essentially the same test-then-split loop) to improve flow.
Best Practices
- Prefer many small, single-purpose tasks over few large multi-purpose ones — easier to test, debug, and swap models per task.
- Give every atomic task a strict, minimal output schema (not free-form prose) so it composes cleanly with downstream steps.
- Keep routing/branching logic in the orchestrating pipeline, never inside an atomic task.
- Name tasks with a verb + object pattern; this doubles as documentation.
- Validate each atomic task in isolation with edge-case inputs before wiring into the pipeline.
- When in doubt, split further — merging two clean atomic tasks later is easier than untangling one messy one.
Common Pitfalls
- Hidden multi-tasking: a task named
analyze-documentthat secretly does classification + extraction + scoring. Split it. - Ambiguous output shape: outputs like "a helpful summary" instead of a defined schema — model behavior becomes unpredictable.
- Conditional logic baked into the task: "if it's a complaint, do X, otherwise do Y" belongs in the pipeline router, not the atomic unit.
- Over-fragmentation: splitting so finely that tasks lose meaningful context (e.g., separating "read subject" from "read body" when both are needed for the same decision). Atomic ≠ trivial — it means single-responsibility, not single-token.
- Skipping isolated validation: assembling the full pipeline before testing each task alone makes failures hard to localize.