AI Skill Report Card

Designing Formative Program Evaluations

B+79·Sep 26, 2026·Source: Extension-page

Designing Formative Program Evaluations for Multidisciplinary Fields

12 / 15

A formative evaluation serves three audiences at once: program staff (who need to adapt practice), funders/policymakers (who need accountability evidence), and researchers (who need generalizable knowledge). Design the evaluation so a single data collection effort feeds all three.

Minimal design template:

  1. Monitoring data (routinely collected by program staff: enrollment, attendance, milestones, exit status)
  2. Participant surveys (pre/post or repeated measures on psychosocial variables, satisfaction, self-efficacy)
  3. Case file / administrative data (structured extraction of qualitative case notes into codeable variables)
  4. Triangulate all three at regular intervals (e.g., quarterly) and feed results back into a dialogical loop with stakeholders (staff, funders, participants where possible).
Recommendation▾
Add a third, more contrasting example (e.g., a low-resource program with limited data infrastructure) to show adaptability beyond the labor-market-integration domain repeatedly used.
14 / 15

Progress:

  • Step 1: Map the multidisciplinary field and stakeholders
  • Step 2: Define formative (not just summative) evaluation questions
  • Step 3: Design multi-source data collection
  • Step 4: Build feedback loops into the program cycle
  • Step 5: Analyze for dual purpose (practice + research)
  • Step 6: Disseminate to differentiated audiences

Step 1: Map the field and stakeholders

Identify which disciplines claim a stake in the program's success (e.g., labor economics, social work, political science, public health). Each discipline will value different outcome types — note this explicitly, since it shapes what "success" data you must collect (e.g., labor market metrics AND psychosocial wellbeing AND service-delivery fidelity).

Step 2: Define formative evaluation questions

Formative questions are about improving the program while it runs, distinct from summative questions about final effectiveness. Frame questions such as:

  • What barriers are participants encountering at each program stage?
  • Which subgroups are underserved or dropping out, and why?
  • Are intended mechanisms of change (e.g., skill-building, social capital, language acquisition) actually occurring?

Step 3: Design multi-source data collection

Use at least three complementary sources so weaknesses of one are offset by another:

  • Monitoring/administrative data: continuous, low-burden, tracks flow and outcomes over time
  • Structured surveys: standardized instruments for psychosocial/attitudinal constructs, repeated at intervals
  • Case file/qualitative data: captures individualized context, converted into codeable categories for mixed-methods integration

Track cohort entry over the full evaluation period (not a single snapshot) to detect trends and allow N to accumulate for subgroup analysis (e.g., 234 participants over 4 years, entering 2018–2022).

Step 4: Build feedback loops into the program cycle

Schedule regular (e.g., annual or semi-annual) reporting moments where preliminary findings are presented back to program staff and funders before the evaluation concludes. This is the defining feature of formative (vs. summative) evaluation — data usefulness comes from timely dialogue, not just a final report.

Step 5: Analyze for dual purpose

Run analyses that simultaneously answer:

  • Practice question: What should the program change now?
  • Research question: What does this reveal about the broader mechanism (e.g., how psychosocial burden interacts with labor market entry for a vulnerable subgroup)?

Report both quantitative outcome patterns (e.g., employment rates, program completion) and qualitative process insights (e.g., how caseworkers adapt practice for specific barriers).

Step 6: Disseminate to differentiated audiences

Produce at minimum:

  • An internal practice brief (actionable, program-specific)
  • A funder/policy report (outcomes tied to program goals and target-population indicators)
  • A peer-reviewed or edited-volume contribution (situates findings in the multidisciplinary literature)
Recommendation▾
Include a concrete template/table for the multi-indicator framework (funder metrics vs. social work metrics vs. researcher metrics) rather than describing it prose-only.
14 / 20

Example 1: Input: A 4-year labor market integration program for migrant women shows high enrollment but unclear whether participants attain skilled employment; funders want proof of effectiveness while staff want to know why some participants disengage early. Output: Combine monitoring data (entry/exit dates, employment status at exit) with repeated psychosocial surveys (self-efficacy, stress) and case file coding of caseworker notes (barriers cited: childcare, language, discrimination). Present quarterly dialogical sessions with staff highlighting early-disengagement risk factors (e.g., low German proficiency + childcare burden), while annual report to funder shows placement rates and skill-shortage-sector matches. Draft a manuscript on psychosocial burden as a predictor of program engagement for the cross-disciplinary literature (social work + labor economics).

Example 2: Input: A multidisciplinary steering committee (economists, social workers, policymakers) disagrees on what "success" means for a new integration pilot. Output: Run a stakeholder mapping session to surface each discipline's implicit outcome definition (economists: wage/employment; social workers: wellbeing/agency; policymakers: cost-per-placement). Design evaluation to capture indicators for all three, and use the formative reporting cycle to negotiate a shared, weighted definition of success documented for the final evaluation report.

Recommendation▾
Show a 'bad outcome' example explicitly (e.g., an evaluation that failed by waiting until program end) to sharpen the contrast between good and poor formative evaluation design.
  • Treat evaluation data collection as a relationship-building tool with program staff, not a compliance exercise — staff who see timely useful feedback provide better-quality data.
  • Always disaggregate outcomes by relevant subgroup (e.g., migration status, education level, family situation) since aggregate "success" rates can mask that a vulnerable subgroup is being poorly served.
  • Use case file/qualitative data to explain why quantitative patterns occur, not just to illustrate them after the fact.
  • Preserve a longitudinal cohort structure (track entry cohort over years) rather than only cross-sectional snapshots, to detect real program evolution vs. noise.
  • Explicitly name the multidisciplinary tension in written outputs — reviewers/readers from each field will expect their own outcome language to be addressed.
  • Don't wait until program end to share findings — this converts a formative evaluation into a summative one and loses the dialogical value that makes it useful to practice.
  • Don't rely on a single data source (e.g., only monitoring data) — administrative data alone cannot explain mechanisms or lived barriers.
  • Don't let funder-defined outcome metrics silently override social-work-relevant wellbeing indicators — negotiate a multi-indicator framework explicitly at the start.
  • Don't treat qualitative case file data as merely illustrative; systematically code it so it can be integrated with quantitative analysis.
  • Don't assume subgroup homogeneity within a broad population label (e.g., "women with migration background") — heterogeneity by origin, education, and family situation drives differential outcomes.
0
Grade B+AI Skill Framework
Scorecard
Criteria Breakdown
Quick Start
12/15
Workflow
14/15
Examples
14/20
Completeness
17/20
Format
14/15
Conciseness
12/15