Designing Formative Program Evaluations
Designing Formative Program Evaluations for Multidisciplinary Fields
A formative evaluation serves three audiences at once: program staff (who need to adapt practice), funders/policymakers (who need accountability evidence), and researchers (who need generalizable knowledge). Design the evaluation so a single data collection effort feeds all three.
Minimal design template:
- Monitoring data (routinely collected by program staff: enrollment, attendance, milestones, exit status)
- Participant surveys (pre/post or repeated measures on psychosocial variables, satisfaction, self-efficacy)
- Case file / administrative data (structured extraction of qualitative case notes into codeable variables)
- Triangulate all three at regular intervals (e.g., quarterly) and feed results back into a dialogical loop with stakeholders (staff, funders, participants where possible).
Progress:
- Step 1: Map the multidisciplinary field and stakeholders
- Step 2: Define formative (not just summative) evaluation questions
- Step 3: Design multi-source data collection
- Step 4: Build feedback loops into the program cycle
- Step 5: Analyze for dual purpose (practice + research)
- Step 6: Disseminate to differentiated audiences
Step 1: Map the field and stakeholders
Identify which disciplines claim a stake in the program's success (e.g., labor economics, social work, political science, public health). Each discipline will value different outcome types — note this explicitly, since it shapes what "success" data you must collect (e.g., labor market metrics AND psychosocial wellbeing AND service-delivery fidelity).
Step 2: Define formative evaluation questions
Formative questions are about improving the program while it runs, distinct from summative questions about final effectiveness. Frame questions such as:
- What barriers are participants encountering at each program stage?
- Which subgroups are underserved or dropping out, and why?
- Are intended mechanisms of change (e.g., skill-building, social capital, language acquisition) actually occurring?
Step 3: Design multi-source data collection
Use at least three complementary sources so weaknesses of one are offset by another:
- Monitoring/administrative data: continuous, low-burden, tracks flow and outcomes over time
- Structured surveys: standardized instruments for psychosocial/attitudinal constructs, repeated at intervals
- Case file/qualitative data: captures individualized context, converted into codeable categories for mixed-methods integration
Track cohort entry over the full evaluation period (not a single snapshot) to detect trends and allow N to accumulate for subgroup analysis (e.g., 234 participants over 4 years, entering 2018–2022).
Step 4: Build feedback loops into the program cycle
Schedule regular (e.g., annual or semi-annual) reporting moments where preliminary findings are presented back to program staff and funders before the evaluation concludes. This is the defining feature of formative (vs. summative) evaluation — data usefulness comes from timely dialogue, not just a final report.
Step 5: Analyze for dual purpose
Run analyses that simultaneously answer:
- Practice question: What should the program change now?
- Research question: What does this reveal about the broader mechanism (e.g., how psychosocial burden interacts with labor market entry for a vulnerable subgroup)?
Report both quantitative outcome patterns (e.g., employment rates, program completion) and qualitative process insights (e.g., how caseworkers adapt practice for specific barriers).
Step 6: Disseminate to differentiated audiences
Produce at minimum:
- An internal practice brief (actionable, program-specific)
- A funder/policy report (outcomes tied to program goals and target-population indicators)
- A peer-reviewed or edited-volume contribution (situates findings in the multidisciplinary literature)
Example 1: Input: A 4-year labor market integration program for migrant women shows high enrollment but unclear whether participants attain skilled employment; funders want proof of effectiveness while staff want to know why some participants disengage early. Output: Combine monitoring data (entry/exit dates, employment status at exit) with repeated psychosocial surveys (self-efficacy, stress) and case file coding of caseworker notes (barriers cited: childcare, language, discrimination). Present quarterly dialogical sessions with staff highlighting early-disengagement risk factors (e.g., low German proficiency + childcare burden), while annual report to funder shows placement rates and skill-shortage-sector matches. Draft a manuscript on psychosocial burden as a predictor of program engagement for the cross-disciplinary literature (social work + labor economics).
Example 2: Input: A multidisciplinary steering committee (economists, social workers, policymakers) disagrees on what "success" means for a new integration pilot. Output: Run a stakeholder mapping session to surface each discipline's implicit outcome definition (economists: wage/employment; social workers: wellbeing/agency; policymakers: cost-per-placement). Design evaluation to capture indicators for all three, and use the formative reporting cycle to negotiate a shared, weighted definition of success documented for the final evaluation report.
- Treat evaluation data collection as a relationship-building tool with program staff, not a compliance exercise — staff who see timely useful feedback provide better-quality data.
- Always disaggregate outcomes by relevant subgroup (e.g., migration status, education level, family situation) since aggregate "success" rates can mask that a vulnerable subgroup is being poorly served.
- Use case file/qualitative data to explain why quantitative patterns occur, not just to illustrate them after the fact.
- Preserve a longitudinal cohort structure (track entry cohort over years) rather than only cross-sectional snapshots, to detect real program evolution vs. noise.
- Explicitly name the multidisciplinary tension in written outputs — reviewers/readers from each field will expect their own outcome language to be addressed.
- Don't wait until program end to share findings — this converts a formative evaluation into a summative one and loses the dialogical value that makes it useful to practice.
- Don't rely on a single data source (e.g., only monitoring data) — administrative data alone cannot explain mechanisms or lived barriers.
- Don't let funder-defined outcome metrics silently override social-work-relevant wellbeing indicators — negotiate a multi-indicator framework explicitly at the start.
- Don't treat qualitative case file data as merely illustrative; systematically code it so it can be integrated with quantitative analysis.
- Don't assume subgroup homogeneity within a broad population label (e.g., "women with migration background") — heterogeneity by origin, education, and family situation drives differential outcomes.