Research Notebook System
Keep the reasoning with the experiments
Research Notebook connects questions, experimental evidence, and the decisions that follow. It keeps this account in linked Markdown files alongside the project's code.
You can adopt the notebook a piece at a time, without learning its layout first. Describe what you want to your coding agent in ordinary sentences, and it creates and updates the files. Setup creates a few core files; plans, spend tracking, and manuscripts are added the first time you ask for them. The guide and reference describe what the agent writes, for when you want to check it.
- Continuity across people and sessions
- Researchers and agents can resume from recorded predictions, results, and next actions.
- Claims traceable to evidence
- A claim identifies the experiment, measured quantity, and design choices that support its stated scope.
An example conversation with the agent
Five requests take the example project from setup to a supported claim. Its job IDs and results are illustrative.
-
YouSet up a research notebook. Our question is whether gradient accumulation reproduces true large-batch training. Jobs run through Slurm.
AgentCreates
lab-notebook/and records the question as RQ1. -
YouDesign a cheap pilot for RQ1. Don't run it yet.
AgentWrites the hypothesis, a prediction that the two losses differ by less than 0.02 at one seed, and a decision rule, then asks you to review them.
-
YouLooks good. Run it.
AgentSubmits Slurm job 48152. When it finishes, records a difference of 0.004 and applies the decision rule: proceed to the three-seed comparison.
-
You, in a new sessionWhere are we? Continue.
AgentReads the status file, finds the pilot result and its next step, and drafts EXP-002, a three-seed comparison, for your review.
-
YouApproved. Run it, then write up what we found.
AgentAfter the run, writes a finding that combines both experiments and adds claim C1, citing
EXP-001:E1,EXP-002:E1, andEXP-002:E2.
Each file name opens that file in the explorer below, except the pilot record, which opens on GitHub. Install the skill to start a notebook in your own project.
Notebook files
Explore the notebook folder, then open an experiment or plan file to see its sections. The two files come from the example notebook, which tests whether gradient accumulation reproduces true large-batch training.
plans/
experiments/
findings/
papers/
reports/
references/
STATUS.md
The project's current state: what's established, what's blocked, and what to do next. Links point to the supporting records.
QUESTIONS.md
The questions the project is investigating, with stable IDs, current status, and links to the experiments or findings that answer them.
PRIORITIES.md
The current research focus and a short queue of next actions. Completed work leaves the queue; experiment records hold the results, and CHANGELOG.md records what was learned or decided.
CHANGELOG.md
A dated history of research results, decisions, and publication milestones. Each entry links to the record that holds the evidence and states what changed for the project.
plans/completed/2026-08-12-accumulation-controls.md
A research objective divided into phases, with limits, review gates, current state, and conditions for completion. The plan coordinates the work and links to the experiments that supply its evidence. Active plans sit directly in plans/; this one is completed.
plans/spend/LEDGER.md
One row per authorized attempt, with the money committed and the money actually spent. A late or wrong figure gets a new correction row; existing rows are never edited, so incurred, committed, and remaining amounts can be totaled from the ledger.
experiments/EXP-002-accumulation-comparison.md
The design and results of one experiment, including what was predicted in advance, what informed the design, observed outcomes, and references to the run artifacts.
findings/2026-08-16-accumulation-matches-large-batch.md
A dated synthesis that draws on several experiments and states the scope of its conclusion and threats to validity.
CLAIMS.md
Claims intended for a paper, linked to specific experiment estimands, with their support status, tested scope, and remaining blockers.
PUBLICATION.md
Publication status for each manuscript, with links to drafts and submission packets, target venues, and blockers to submission.
GLOSSARY.md
Definitions of project-specific terms, symbols, and acronyms, with links to the evidence or method decisions behind them.
BIBLIOGRAPHY.md
An annotated bibliography explaining what each cited paper says and why it matters to the project. Links lead to source papers and any local copies or notes in references/.
papers/name.typ
A manuscript authored by the project. Its claims draw on the notebook evidence; cited third-party papers belong in references/.
reports/analysis.md
A living analysis with room for figures and detail beyond a dated finding. It can change as the project accumulates evidence.
references/source-notes.md
Notes on a cited source, kept with archived third-party papers and their provenance. Manuscripts authored by the project belong in papers/.
Fixed before results
Recorded as decisions are made
Recorded during and after the run
Title and header fields
The experiment ID and title, the created date, the lifecycle status, and the research questions it addresses. Status says where the experiment is in its lifecycle; it does not record approval.
# EXP-002: Accumulation comparison
**Created**: 2026-08-12
**Status**: completed
**Research questions**: [[RQ1]] Hypothesis
A specific, falsifiable claim the experiment tests.
## Hypothesis
Across seeds 1 through 3, the mean final-validation-loss difference between
accumulated and true batches is below 0.02. Method
The instrument, conditions, seeds, metric, code revision, and the command that produces the result, so that another session can rerun it.
## Method
- **Instrument**: deterministic `scripts/simulate.py` toy simulator
- **Conditions**: `true-batch` and `accumulated`
- **Seeds**: 1, 2, and 3
- **Primary metric**: mean paired final-validation-loss difference
- **Revision**: `synthetic-code-v1`
- **Key command**: `python3 scripts/simulate.py --condition <condition> --seed <seed>` Estimands
The quantities the experiment measures, one ### E# heading each. Each heading carries a registration value fixed before results, which separates planned tests from discoveries made in the same run.
## Estimands
### E1 Mean paired loss difference
**Registration**: registered. The statistic is the mean paired
… E1 Mean paired loss difference
One registered quantity with its statistic and threshold. A claim cites it by ID, as EXP-002:E1, so a reader can trace the claim to this definition and its result.
### E1 Mean paired loss difference
**Registration**: registered. The statistic is the mean paired
final-validation-loss difference across seeds 1 through 3; the threshold is
0.02; a mean at or above it fails P1. E2 Per-seed loss difference
A second registered quantity: every seed must pass, not only the mean. Because it is registered, its result counts as a planned test.
### E2 Per-seed loss difference
**Registration**: registered. Every paired difference must be below 0.02; one
seed at or above it fails P2. Informed by
What was read or measured before the design was fixed. Here the pilot already used seed 1, so E1 is a replication at that seed and a fresh test only at seeds 2 and 3.
## Informed by
- [[EXP-001-accumulation-pilot]], whose registered E1 passed at seed 1. Seed 1
is reused here, so E1 in this record is a replication at that seed and a
fresh test at seeds 2 and 3.
- [[2026-08-12-accumulation-controls]] Preregistered predictions
The predicted outcome for each estimand, written before the run so the result can be scored against it.
## Preregistered predictions (a priori)
- **P1: Mean equivalence (E1)**: the mean paired difference is below 0.02.
- **P2: Seed consistency (E2)**: every paired difference is below 0.02. Decision rule
What to do for each material outcome, decided in advance. The session that processes the result applies the rule instead of deciding afresh after seeing the numbers.
## Decision rule (a priori)
- **If P1 and P2 hold**: support a synthetic-scale equivalence finding.
- **Otherwise**: keep RQ1 open and report the failing seeds. Decisions
Dated design choices between credible options. Each entry names the option that lost, what the choice costs, and what settled it. Entries are added and never edited.
## Decisions
- **2026-08-14** — Kept the 0.02 margin rather than tightening it to 0.01 after
[[EXP-001-accumulation-pilot]] came in at 0.004: a margin chosen after seeing
the pilot would depend on the result it is meant to judge. Costs sensitivity
to differences between 0.01 and 0.02. Source: pilot-gate review.
- **2026-08-14** — Reused seed 1 rather than drawing seeds 4 to 6: the paired
comparison stays anchored to the pilot. Costs one fresh seed; E1 is a
replication at seed 1 and a fresh test only at seeds 2 and 3. Source:
pilot-gate review. Human review
The approvals at each checkpoint: the design, the predictions, and the interpretation. A status value does not record approval; this section does.
## Human review
- **Design**: paired-seed comparison approved after the pilot gate.
- **Preregistration**: predictions and threshold approved before the run.
- **Analysis and interpretation**: results, limitations, and proposed finding
reviewed before synthesis. Runs
Each execution attempt: backend, job ID, description, status, and where its artifacts are. Usage in units belongs with the attempt; money goes in the spend ledger.
## Runs
| Backend | Job ID | Description | Status | Artifacts |
|---|---|---|---|---|
| Slurm | 48161 | Both conditions at seeds 1 through 3 | completed | `results/EXP-002/48161/metrics.json` | Results
Observed values with their uncertainty and effect sizes, and the checks that the output is valid.
## Results
| Seed | True batch | Accumulated | Paired difference |
|---|---:|---:|---:|
| 1 | 0.799 | 0.803 | 0.004 |
| 2 | 0.803 | 0.807 | 0.004 |
| 3 | 0.796 | 0.800 | 0.004 |
The mean paired difference was 0.004.
… Outcomes against preregistered predictions
Each prediction scored as confirmed, partial, refuted, null-confirmed, or not tested, beside what was predicted and what was observed.
### Outcomes against preregistered predictions
| Prediction | Verdict | Predicted | Observed |
|---|---|---|---|
| P1 | confirmed | Mean difference < 0.02 | 0.004 |
| P2 | confirmed | Every difference < 0.02 | All three were 0.004 | Conclusion
What the result means, the scope at which it was tested, and threats to validity.
## Conclusion
Both preregistered equivalence checks pass within the deterministic simulator.
The design does not test optimizer state, floating-point order, or a real model. Follow-ups
Concrete next steps the result calls for, checked off when done.
## Follow-ups
- [x] Synthesize the pilot and comparison in a finding. Artifacts
Where the analyzed output lives and the revision that produced it. A full record also gives a content hash and retrieval date.
## Artifacts
- `results/EXP-002/48161/metrics.json`, fictional artifact at revision `synthetic-code-v1` Findings
Links to the cross-experiment findings that use this result.
## Findings
- [[2026-08-16-accumulation-matches-large-batch]] Kept current throughout
Set when the plan is approved
Recorded as the plan runs
Added at close
Frontmatter
Status, a one-line summary, the single next action, owner and reviewer, current phase, and dates. The status, current phase, and next action change as the plan moves, so a new session knows where to resume.
---
status: completed
summary: Test synthetic gradient accumulation against true large batches
next_action: none
owner: example maintainer
reviewer: example reader
current_phase: Closed
created: 2026-08-12
updated: 2026-08-16
--- Objective
The question the plan answers and the result that would settle it.
## Objective
Determine whether the synthetic accumulated condition stays within 0.02 final
validation loss of the true-batch condition. Existing evidence
What was already known when the plan was written, so a later reader can tell which evidence the plan itself produced.
## Existing evidence
RQ1 was open when this plan was created. No result had been observed. Phases
The work in order. Each phase names its experiments and the condition for moving on. A larger plan divides phases into work packets with explicit limits.
## Phases
### Phase 1: Cheap validation
Run seed 1 in both conditions. Continue only if both results are finite and the
paired difference is below 0.02. Record [[EXP-001]].
… Phase 1: Cheap validation
A cheap check that gates the expensive one: continue only if the pilot result is finite and inside the margin.
### Phase 1: Cheap validation
Run seed 1 in both conditions. Continue only if both results are finite and the
paired difference is below 0.02. Record [[EXP-001]]. Phase 2: Three-seed comparison
The full comparison, run only after the pilot gate passes and the evidence is reviewed. Its success condition matches the experiment's registered estimands.
### Phase 2: Three-seed comparison
Run seeds 1 through 3 in both conditions. Support the scoped claim only if the
mean and every paired difference are below 0.02. Record [[EXP-002]]. Risks and controls
What could make the result misleading, and the control the plan applies to each risk.
## Risks and controls
- The simulator is not a training system; keep every conclusion synthetic.
- Write predictions and gates before recording outputs.
- Use separate job IDs and ledger records for the two phases. Terminal conditions
When the plan is complete, blocked, or abandoned, decided before execution begins.
## Terminal conditions
- Complete when both phases pass and their evidence is durable.
- Block if an output is non-finite or an artifact cannot be verified.
- Abandon if the simulator cannot represent both conditions deterministically. Human review
The reviews at each gate: the plan before execution, each phase's evidence before the next phase, and the disposition before closure. Approving a plan does not give an agent additional permissions.
## Human review
- The plan and Phase 1 design were approved before execution.
- Phase 1 evidence was reviewed at the pilot gate before Phase 2 was approved.
- Phase 2 evidence and the terminal disposition were reviewed before closure. Decisions
Dated choices between credible options, each naming the option that lost, what the choice costs or forecloses, and what settled it. Entries are added and never edited.
## Decisions
- **2026-08-16** — Closed the plan as completed rather than adding a
real-training phase after [[EXP-002-accumulation-comparison]] passed: the
objective was scoped to the simulator. Forecloses any claim about real models
from this plan; a replication would be a new plan. Source: closure review. Completion report
Each goal marked met or not met, with the record that shows it, plus limitations and open follow-ups. A superseded or abandoned plan writes a Disposition instead.
## Completion report
Completed on 2026-08-16. Both phases passed their preregistered gates.
| Goal | Outcome | Record |
|---|---|---|
| Phase 1 pilot is finite and inside the margin | met | [[EXP-001-accumulation-pilot]] |
… Evidence
Every record the plan's conclusion rests on: experiments, findings, and processed job records.
## Evidence
- [[EXP-001-accumulation-pilot]]
- [[EXP-002-accumulation-comparison]]
- [[2026-08-16-accumulation-matches-large-batch]]
- Ledger records `jobs/processed/slurm/48152.json` and `48161.json` This plan is small enough to need no resource estimate or confirmation reserve. A plan that uses shared compute or paid services states its resource estimate in units, and a plan meant to produce a confirmatory claim adds a confirmation reserve.
Indexes and dashboards draw from these files. Add an optional record when the project needs it.
System features
| Feature | Benefit |
|---|---|
| Portable, version-controlled files | Markdown files work offline, support Git/jj diffs, and are searchable with ordinary tools. Obsidian provides linked navigation. The notebook can have its own version-control repository. |
| Research questions organize the work | Each experiment links to its research questions and keeps its hypothesis, method, runs, results, and interpretation together. Findings combine evidence from several experiments. |
| One home for each kind of information | STATUS describes the project's current state; QUESTIONS tracks inquiry; PRIORITIES identifies next work; PUBLICATION tracks submission readiness. Each links to the supporting records, so a result can be updated in one place. |
| Handoffs between sessions and agents | One session can design and launch an experiment, another process its results, and a third continue the work. Recorded predictions, decision rules, and follow-ups let each recover the reasoning behind the next step. |
| Registration for each measured quantity | An estimand is the quantity an experiment seeks to estimate. The notebook records what was fixed in advance for each estimand, plus what earlier evidence informed the design. This distinguishes planned tests from discoveries within the same run and makes post-hoc reinterpretation easier to detect. |
| Traceable claims and provenance | A claim cites a specific experiment estimand, such as EXP-042:E2. Links connect plans, experiments, jobs, claims, and papers. A reader can trace a claim's support and identify what needs reconsideration when evidence changes. |
| A result-processing workflow | After a run finishes, validate the output and instrument, record the result, apply the decision rule, update dependencies, preserve the artifacts, and report. Process failed and null results too, so their evidence and follow-ups remain available. |
| Cost estimation and spend tracking | Plans estimate work in units that stay reproducible: GPU-hours, API calls, tokens, elapsed time. Experiment records log usage for each attempt. All money is kept in plans/spend/: the owner's approved ceiling in AUTHORITY.json, and an append-only ledger of what each attempt committed and cost. Incurred spend, committed spend, and remaining headroom are calculated from the ledger. A run gets its budget as seconds or calls, never dollars, and the validator flags any currency amount in an experiment record. |
| Mechanical checks and review tracking | Validators check defined formats and references. Human review decisions stay with the experiments, plans, and claims they govern. Content-hash review tracking can identify changed artifacts that need another review. |
| A path from evidence to publication | CLAIMS.md links paper claims to their evidence; PUBLICATION.md records paper scope, venues, and blockers. Manuscripts and supporting notes stay linked to the experiments they draw on. |
Session handoffs
- Design and launch
Write the hypothesis, predictions, and decision branches before inspecting results. Record what earlier evidence informed the design. Estimate the work in units and get its spend approved.
- Process the result
Check the output and instrument, record the outcome, apply the decision rule, update dependencies, and preserve the artifacts.
- Apply the recorded decision
The next session follows the decision rule and follow-ups. Failed and null results can close a branch as well as open one.
Check, interpret, and record results before moving on to the next experiment. The same requirement applies to laptop runs and remote jobs.
Installation
The public agent skill supplies the notebook conventions and workflows.
1. Install in a terminal
Run this command from your research project's directory:
npx skills add osteele/research-notebook -s research-lab-notebook -y 2. Ask your coding agent to set up the notebook
After installation, send this prompt to the agent working in your project:
Use $research-lab-notebook to add a research notebook to this project. Jobs run through Slurm.
Name your project's runner or explain how you run scripts locally. The skill supplies instructions for working with those tools; it does not install a scheduler or research loop.
Start with the records your project needs, and keep them current as the research changes.
For designing a follow-up or assessing a claim, use the evidence and data-reuse workflow. It covers experiment records, confirmation reserves, and the claims ledger.
Integration with other research tools
Experiment files connect hypotheses and interpretations to outputs from metric trackers, computational notebooks, version control, and job runners. The spend ledger connects each attempt to the provider's charges for it.
| Tool | What it manages | Connection to the notebook |
|---|---|---|
| Weights & Biases / MLflow | Metrics, run comparisons, and hyperparameters. | Experiment records link to runs and recorded metrics, then document the hypothesis, interpretation, and resulting claims. |
| Jupyter | Interactive exploration and computational analysis. | Experiment records reference computational notebooks and their outputs, preserving the design and conclusions alongside the analysis. |
| Git / jj | Code history and versioned changes. | Notebook files are versioned; experiment records identify the code revision that produced a result. |
| Slurm / SkyPilot / Weft | Job execution, status, logs, and output retrieval. | Run records retain job IDs and artifact locations. Result processing records whether the output was validated, interpreted, and acted on. |
| Provider billing | Charges for compute and API calls, from cloud consoles and API usage dashboards. | Spend-ledger rows are keyed to job IDs. When a final bill differs from the recorded charge, the difference goes in as a correction row in plans/spend/LEDGER.md. |
These connections use file links and project-specific commands. API integration requires a separate connector.
Read and edit the files in Obsidian, a text editor, or a coding agent.