Building Reproducible Jupyter Workflows
For any new computational research project, set up this structure before writing analysis code:
project/
├── environment.yml # or requirements.txt / Pipfile
├── README.md # how to reproduce, in 3 steps or fewer
├── binder/ # Binder config for one-click reproduction
├── data/
│ ├── raw/ # never modified, read-only
│ └── processed/ # generated by notebooks, gitignored
├── notebooks/
│ ├── 01-data-preparation.ipynb
│ ├── 02-analysis.ipynb
│ └── 03-figures.ipynb
├── src/ # reusable code extracted from notebooks
└── tests/ # tests for src/, run in CI
Pin the environment immediately:
Bashconda env export --no-builds > environment.yml # or pip freeze > requirements.txt
Then verify reproducibility from a clean environment before continuing:
Bashconda env create -f environment.yml -n repro-check conda activate repro-check jupyter nbconvert --to notebook --execute notebooks/*.ipynb
Progress:
- Step 1: Define the environment specification (exact package versions, not ranges)
- Step 2: Organize notebooks by pipeline stage, numbered in execution order
- Step 3: Extract reusable logic from notebooks into a tested
src/package - Step 4: Ensure notebooks execute top-to-bottom with no hidden state (restart & run all)
- Step 5: Separate data (raw/immutable) from generated outputs
- Step 6: Add a Binder/Docker config so others can run it with zero local setup
- Step 7: Strip or version-control notebook outputs sensibly (nbstripout for git)
- Step 8: Archive a DOI-citable snapshot (Zenodo + GitHub release) at publication time
- Step 9: Write a README with exact reproduction steps, tested by someone else
Step 1 — Environment specification. Never rely on "whatever is installed." Pin exact versions of Python, Jupyter, kernels, and all dependencies. Record OS if it matters (e.g., compiled extensions).
Step 2 — Notebook organization. One notebook = one stage of the pipeline (ingest, clean, model, visualize). Avoid monolithic notebooks mixing exploration and final analysis — keep exploratory work in a separate scratch/ folder excluded from the reproducible path.
Step 3 — Extract to src/. Any function used more than once, or any function requiring testing/debugging, moves out of the notebook into an importable module. Notebooks should read like a narrative calling well-tested functions, not contain the implementation.
Step 4 — No hidden state. Before considering a notebook "done," restart the kernel and run all cells in order. Out-of-order execution is the single most common reproducibility failure.
Step 5 — Data separation. Raw data is read-only and never edited in place. All derived data is regenerable from raw data + code — never hand-edited.
Step 6 — Zero-setup reproduction. Provide a binder/ directory (or Dockerfile) so a reviewer or collaborator can click a link and get a running, identical environment without local installation.
Step 7 — Output handling in git. Decide deliberately: either strip all outputs before commit (clean diffs, use nbstripout) or commit outputs intentionally (for rendered documentation) — never commit accidentally stale outputs.
Step 8 — Archival snapshot. At submission/publication, tag a release and mint a DOI via Zenodo (or similar) so the exact code state is permanently citable, independent of future repository changes.
Step 9 — README test. The README's reproduction instructions must be validated by someone other than the author, ideally on a different machine.
Example 1:
Input: A researcher has a single 2000-line notebook mixing data download, cleaning, model fitting, and all figures, with no environment file, that "works on my machine."
Output: Split into 01-download-data.ipynb, 02-clean-data.ipynb, 03-fit-model.ipynb, 04-generate-figures.ipynb. Move the model-fitting function and cleaning routines into src/models.py and src/cleaning.py with unit tests. Add environment.yml pinned from conda env export. Add a Binder config. Verify full pipeline reruns cleanly in a fresh conda env before calling it done.
Example 2:
Input: Paper accepted, need to make the computational experiments citable and permanently reproducible.
Output: Tag a GitHub release of the final notebook state, connect the repo to Zenodo, publish the release to mint a DOI, and cite that DOI (not the GitHub URL) in the paper. Confirm the Zenodo archive includes environment.yml and that jupyter nbconvert --execute succeeds against the archived snapshot in a clean container.
- Pin exact dependency versions; "latest" is not reproducible.
- Prefer narrative notebooks calling tested library code over notebooks containing all logic inline.
- Always test "restart kernel, run all" before sharing or committing.
- Separate raw, immutable data from derived, regenerable data.
- Make reproduction a single command (
make reproduce, or onenbconvertcall), not a sequence of undocumented manual steps. - Cite a DOI-archived snapshot (Zenodo), not a mutable GitHub URL, in publications.
- Treat notebooks as the reproducible entry point and
src/as the tested, reusable foundation.
- Running cells out of order and getting correct results only because of leftover in-memory state.
- Committing large data files or notebook outputs without deciding on a policy, bloating git history.
- Specifying dependencies with loose version ranges that silently break months later.
- One giant notebook with no separation between exploration, final analysis, and reusable code.
- Forgetting to test the environment/reproduction steps on a machine other than the author's own.
- Citing a live GitHub repo link instead of an immutable, versioned archive (Zenodo DOI) in publications.