Review Ledger
REVIEW-LEDGER.md is an optional notebook file that records independent reviews and
turns the faults they find into gates: questions to ask before the next
experiment is written or submitted. Each gate exists because a run failed that way. Reviews
are requested by a person or required by project instructions at a named transition; the
ledger records them and does not schedule them.
The independent review guide covers when to request a review, reviewer independence, and a worked example.
review ledger / layout · source
File layout
# Review ledger
## Review register
| ID | Date | Kind | Target | Reviewer | Relation | Findings |
| --- | --- | --- | --- | --- | --- | --- |
| R001 | 2026-09-10 | review-script | scripts/exp_002_comparison.py @ 3f1c9a2e | fresh session | same model, no shared context | F001, F002 |
| R002 | 2026-09-11 | review-script | scripts/exp_002_comparison.py @ 8b04d17c | second provider | different model | none |
## Open findings
| ID | Fault | Caught by | Severity | Status |
| --- | --- | --- | --- | --- |
| F002 | Summary averages over all cells, including the excluded pilot cell | R001 | silent | repair landed; awaiting review |
## Gates — checkable from the artifact
#### G1 — Does the summary aggregate over the intended slice?
**Added** 2026-09-10, at EXP-002. **Severity:** silent. **Tally: 1** — F001 (founding). **Slug:** `aggregation-slice`.
What failed, in two or three sentences, with the experiment it cost.
**Check:** the one question a reviewer answers from the artifact.
## Gates — need project history
## Promoted
The validator requires a ## Review register section when the file
exists, and checks the ## Open findings table's columns when that
section is present. Edit the ledger through a tool that allocates IDs and serializes
concurrent writers when one is available. Without one, append rows by hand and never
renumber or delete them.
review ledger / review register
Review register
Every review gets a row, including clean ones; similar reviews may share a summary row.
Columns, in order:
ID | Date | Kind | Target | Reviewer | Relation | Findings.
| Column | Meaning |
|---|---|
ID | R-prefixed and unique. Many similar reviews may share one summary row, such as R040–R061: 22 script reviews, 2 findings. |
Kind | The kind of review, such as review-script: a plan, experiment design, experiment script, or results review. |
Target | The exact bytes reviewed: a path with a content digest or revision. A review of earlier bytes does not cover a revision. |
Reviewer, Relation | Who reviewed, and how the reviewer relates to the author. |
Findings | The finding IDs the review raised, or none. |
Clean reviews are the denominator: without them, neither the practice's value nor any single
gate's value can be measured. The Relation column takes one of
self, same model, no shared context, different model, human.
| Relation | Meaning |
|---|---|
self | The author reviewed their own work. |
same model, no shared context | A fresh session of the same model, given only the artifact and its record. |
different model | A reviewer from a different model family or provider. |
human | A person other than the author. |
review ledger / open findings
Findings
Columns, in order:
ID | Fault | Caught by | Severity | Status.
Severity is one of
silent or loud:
silent when the fault would have produced a wrong number with no error, and
loud when it would have failed visibly.
Record that a repair landed, and close the finding when a review of the repaired bytes settles it.
review ledger / gates
Gates
Each gate entry has a heading of the form #### G1 — Question?, a metadata line
with the date and experiment where it was added, its severity, its tally with the findings
attached to it, and a slug; two or three sentences on what failed and the experiment it cost;
and a Check line stating the one question a reviewer answers.
Add a gate when a run fails for a reason no existing gate covers. When a later finding is another instance of an existing gate, attach it and increment the tally instead of adding a near duplicate. Resemblance between two faults is not by itself evidence that they are the same.
| Section | Contains | Given to |
|---|---|---|
| Gates checkable from the artifact | Questions decidable from a script and its output | Any reviewer, including a fresh-context one |
| Gates that need project history | Questions such as whether a sibling experiment already solved this problem | Only reviewers who have that history |
A reviewer is never asked to work out which gates it can reach. Before submitting a new or modified experiment script, read the gate titles and open the detail only for gates that plausibly apply.
review ledger / promoted
Promotion
A gate with a high tally is a candidate for promotion: into a template, a shared checklist,
or a mechanical check that makes the fault impossible to write. Record the promotion under
## Promoted and say what it does not cover; a template reaches only records
written from it. A gate that stops firing is a candidate for retirement.