Applying Scientific Cycle to Social Science
Applying the Scientific Cycle to Social Science
Turn a question about people, markets, institutions, or policy into a testable, revisable research argument. Adapt the cycle to the user's study stage and requested deliverable — do not force completed research back through every step. Treat conclusions as conditional on measurement, context, design, and assumptions.
- Identify topic, research stage, claim type, and available evidence. State working assumptions where details are missing.
- Walk through the scientific cycle (observe → hypothesize → predict → test → analyze → revise), making each link explicit.
- Populate the Study Brief template. Mark unknowns explicitly; distinguish proposed work from completed analysis.
- Before asserting causation, check against the bad example and corrected contrast.
Example invocation: "Use the scientific-cycle skill to develop a research plan on whether public transport reliability improves employment."
Minimal filled-out example — hypothetical research plan:
- Question and scope: Does improving bus reliability increase registered employment after six months among unemployed, transit-dependent adults in one city? Claim type: causal; findings pending.
- Observation and hypothesis: Commuting barriers are a proposed concern, not a verified finding. More reliable buses may improve job access; local job growth is a rival explanation.
- Prediction and test: Employment should rise more among affected residents than in a credible comparison group. Use administrative employment records and measured bus reliability before and after a service change; consider difference-in-differences if parallel trends and limited spillovers are defensible.
- Analysis and interpretation: Estimate the employment difference with uncertainty and assess concurrent policies and residential selection. An imprecise estimate would leave the hypothesis unresolved.
- Ethics and next cycle: Use authorized, protected records and document the analysis. Check that reliability actually improved, then test the commuting mechanism with new evidence.
Progress:
- Frame the inquiry (population, setting, unit of analysis, claim type)
- Observe and question (source, verify pattern isn't an artifact of measurement)
- Form hypothesis + rival explanations (operationalize key concepts)
- Deduce predictions (direction, horizon, subgroup, falsifying evidence)
- Design test matched to question (check method-specific assumptions)
- Analyze (effect size + uncertainty, not just significance)
- Revise and plan next discriminating test
1. Frame the inquiry
State population, setting, period, unit of analysis, and question type:
- Descriptive — what happens
- Predictive — what is likely
- Causal — what changes because of an intervention
- Interpretive — how people understand/experience something
Separate empirical claims from normative judgments — evidence informs policy but doesn't supply the values used to rank options.
2. Observe and question
Identify a documented pattern, anomaly, or lived experience. Check whether sampling, measurement, or shifting definitions could produce it. Record the source; distinguish established observations from proposed ones.
3. Form a hypothesis
Propose a mechanism and plausible rival explanations. Specify scope and what observation would count against it. Operationalize abstract concepts (trust, inequality, welfare) and flag construct-validity concerns, especially across groups.
4. Deduce predictions
Derive observable implications from the mechanism plus explicit auxiliary assumptions: outcome, comparison, expected direction, horizon, relevant subgroups. State magnitude only if justified. Prefer predictions that discriminate between the hypothesis and its rivals. Pre-register confirmatory predictions where feasible; label anything else exploratory.
5. Design a test and gather evidence
Match method to question: surveys, interviews, ethnography, archival/administrative records, experiments, natural experiments, formal models. A simulation shows implications under assumptions — not empirical validation. For causal claims, define treatment, outcome, target effect, and counterfactual, then apply the method checks below.
Method-specific assumption checks:
- RCT — attrition, compliance, interference between units
- Diff-in-diff — credible parallel trends, timing, anticipation effects
- RDD — continuity around cutoff, no manipulation of the running variable
- IV — relevance, exclusion restriction, independence — argue, don't just assert
6. Analyze and assess
Report evidence, uncertainty, competing interpretations, limitations. Favor effect sizes with uncertainty over significance thresholds. Address missingness, multiple testing, sampling, dependence. For qualitative work: justify case selection, coding, negative cases, reflexivity. Never fabricate data, estimates, citations, or completed tests.
7. Revise and repeat
State whether evidence supports, challenges, or leaves the hypothesis unresolved — within scope. Distinguish a weak/imprecise test from genuine evidence of no effect. Propose the next study that could distinguish remaining rival explanations.
- Induction — move from cases/patterns to a candidate explanation. Strengthens generalization across settings but never proves a universal rule; weigh representativeness and rival mechanisms.
- Deduction — move from theory + assumptions to a specific prediction. Valid logic doesn't mean the premises hold. A failed prediction may indict the mechanism, the auxiliary assumptions, the measurement, or the design — not necessarily the core theory. A successful prediction doesn't uniquely confirm it if rivals predict the same thing.
- Abduction (when useful) — infer the most plausible explanation among alternatives, then seek evidence that discriminates between them.
- Association ≠ causation. Consider reverse causality, omitted variables, selection, simultaneity, measurement error. Adding controls doesn't fix identification — controls that are mediators or colliders can introduce new bias.
- People react to policy and to being observed. Consider strategic behavior, spillovers, feedback, equilibrium adjustment. State whether an effect is local, short-run, or likely to generalize.
- Context limits generalization. Institutions, culture, history, and sample selection bound external validity. Motivate subgroup analysis from theory, not post-hoc search (apply the confirmatory/exploratory distinction).
- Ethics. For proposed data collection: consent, confidentiality, harm, institutional review. Use anonymized/controlled-access data when public sharing would expose participants.
- Reproducibility vs. replication. Reproducing = same data/code, same result. Replication = new, independent evidence. Don't conflate the two.
Markdown# Study Brief: [Title] **Stage and evidence status:** [Exploration / research plan / completed analysis] **Claim type:** [Descriptive / predictive / causal / interpretive]
[Research question; population, setting, period, unit of analysis; empirical vs. normative separation]
[Motivating pattern; sources; verified vs. unverified premises]
[Mechanism, scope, induction trail, competing explanations]
[Observable implications; direction/horizon/comparison; disconfirming evidence; exploratory vs. confirmatory]
[Operational definitions; data source/collection plan; sample selection; validity, missingness]
[Method and fit; identification argument; treatment/outcome/counterfactual for causal claims]
[Planned/completed analysis; uncertainty; sensitivity checks; competing interpretations]
[Pending, or: observed findings, warranted conclusion, unresolved alternatives, generalization limits]
[Participant protections; data/code/versioning needed to reproduce]
[What to keep/revise/leave open; next discriminating test]
Example 1: Input: "Does minimum wage increase unemployment among teens?" Output: A study brief that (a) frames this as a causal question on a specific population/period, (b) notes the long empirical literature's mixed findings as context rather than settled fact, (c) proposes a diff-in-diff or border-discontinuity design comparing adjacent regions with different minimum wage changes, (d) flags parallel-trends and spillover (cross-border employment) threats, (e) specifies that findings are local to the studied period/labor market and shouldn't be generalized without further replication.
Example 2: Input: "I already ran a survey showing trust in local government rose after a transparency portal launched. Did I prove the portal worked?" Output: Flags that this is pre/post observational, not causal — rival explanations (concurrent local events, response bias, regression to the mean) remain uncontrolled; suggests looking for a comparison municipality without the portal, checking for pre-trends, and reporting the finding as "associated with" rather than "caused by" pending a stronger design.
Example 3 — completed and analyzed study (hypothetical): All figures and results below are invented for illustration, not empirical evidence.
Input: "We completed a randomized trial of job-search coaching with 1,200 unemployed adults in one city. Assignment was individual: 600 received an offer of coaching and 600 received usual services. Our pre-registered primary outcome was registered employment after six months, measured in administrative records available for all participants. Employment was 44% in the offer group and 38% in the comparison group. The intention-to-treat estimate was +6 percentage points (95% confidence interval: +0.5 to +11.5). Assignment checks found no implementation deviations; we detected no sharing of coaching services across groups, although indirect spillovers remain possible. How should we write up the completed study?"
Output: A completed study brief that:
- States the causal question and population, documents random assignment and complete outcome coverage, and identifies the effect as that of offering coaching rather than receiving it.
- Reports the +6 percentage-point estimate and its confidence interval as supplied, distinguishing the point estimate from the range of effects compatible with the analysis. Notes that statistical significance alone does not establish policy value; costs and a substantively meaningful effect threshold also matter.
- Concludes that the trial provides evidence that offering coaching increased registered employment at six months in this sample, conditional on correct randomization and limited interference. Separates this observed result from the untested explanation that coaching improved search skills.
- Explains that incomplete take-up would not invalidate the intention-to-treat comparison but would change its interpretation; does not infer a treatment-on-the-treated effect without additional analysis and assumptions.
- Limits generalization to the studied population, city, outcome, and horizon. Does not infer sustained employment, higher earnings, or employment effects from scaling the program across the labor market.
- Documents consent and authorized administrative linkage, supplies the pre-registration and reproducible analysis materials under appropriate access controls, and reports any protocol deviations. Does not claim independent reproduction or replication unless verified.
- Revises the research position from a plausible hypothesis to evidence supporting a local, six-month effect of the offer. Proposes longer follow-up and replication in another labor market, with new mechanism or subgroup hypotheses