Spend and Units

documentation/ guide and reference

Money has one home in a notebook: plans/spend/. Its AUTHORITY.json holds the ceilings and allocations the owner approved, and its LEDGER.md holds money committed and spent per authorized attempt. Experiment records, plans, and experiment scripts state scientific quantities and engineering units, and carry no currency amount of any kind.

The cost estimation and tracking guide follows a synthetic campaign through estimate, authorization, tracking, and settlement.

spend / surfaces · source

Surfaces

SurfaceCarriesNever carries
Experiment recordScientific quantities; compute quantities such as GPU-hours, device class, forward passes, API calls, tokens, and retained bytes; elapsed time; Backend and Job ID for each attemptCurrency amounts of any kind
PlanScope, packages, the experiments in each package, and the resource estimate in unitsCeilings, allocations, rates, spend figures
Experiment scriptThe same units, and a runtime cap received as seconds, forward passes, calls, or tokensCurrency constants, rates, rate × time arithmetic, currency flags
plans/spend/AUTHORITY.jsonThe ceilings and allocations the owner approved, who approved them, when, and their history—
plans/spend/LEDGER.mdMoney committed and spent per authorized attempt—

Scientific budgets, such as a retry budget, a strike budget, or an alpha budget, are design quantities and stay in the record. The validator reports a currency amount in an experiment record as an error.

A price is set when a job is placed and can change later, so a rate copied into a record goes stale; the unit counts stay fixed. A ceiling copied into a script or record becomes a second home for it, and the copy goes stale: a script that still checks an old figure refuses the correct authorization. Records are published as preregistration artifacts and paper supplements, while rates and an owner's ceilings are private operating detail.

spend / units in records and plans

Units

Keep them separate. Parallel execution shortens elapsed time without reducing GPU-hours; a queue wait lengthens elapsed time without consuming compute. A projection across providers or devices keeps the configuration and drops the rate: "two 8-GPU nodes for 3 hours" stays true when the price changes.

spend / prices as data

Prices as data

The rule governs spend on the research itself. A price can also be the object of study, or an input to the model under study. Those dollars would mean the same thing if someone else had paid for the compute, so they are scientific quantities. A record keeps them between an opening usd: marker and a closing /usd marker:

<!-- usd: measured — simulator objective: total cost per job -->
| Policy | Cost per job (USD) |
|---|---|
| greedy | $1.20 |
<!-- /usd -->
LabelUse
measured Dollars a simulator or model outputs as its objective. They are experimental measurements.
parameter An external market input such as a spot price, a total-cost-of-ownership rate, or a market-derived cost-efficiency figure. The reason names the source or the model it feeds.

The marker wraps lines, not files. One record often holds both a simulator table and a line about what the rental cost, and a file-level exemption would license the second under cover of the first. Nothing is inferred from position, because scientific dollars appear in prose as well as tables. The validator skips marked lines, reports a marker that is unclosed, nested, stray, unlabeled, or missing its reason, and adds a note with each record's marked line count so growth in exemptions stays visible.

The marker covers figures that remain once spend is out of the record, never spend itself. A study of cost-effectiveness reports the units that determine the price and marks the per-unit price as a parameter.

spend / AUTHORITY.json

Authority file

Both spend files are keyed by the plan's filename stem, never its path. Plans move between status directories, and a file stored beside its plan would be orphaned by an ordinary status change. The validator reports an authority entry whose stem names no plan.

{
  "schema": "plan-spend-authority/v1",
  "plans": {
    "2026-09-01-retrieval-comparison": {
      "shared_ceiling_usd": 21.0,
      "authorized_by": "A. Researcher",
      "recorded_utc": "2026-09-09T00:00:00+00:00",
      "allocations": {
        "confirmation": {
          "ceiling_usd": 8.0,
          "experiments": ["EXP-004"],
          "note": "Held for independent confirmation after the comparison gate."
        }
      },
      "history": []
    }
  }
}
FieldMeaning
schemaAlways plan-spend-authority/v1.
shared_ceiling_usdThe plan's single ceiling, inclusive of everything already spent. Version 1 records amounts in US dollars.
authorized_by, recorded_utcWho approved the current ceiling, and when.
allocationsOptional sub-ceilings inside the shared ceiling, which do not add to it. Each names a ceiling_usd, its experiments, and a note. A ceiling_usd of null names a package whose terms are in its note without a separate figure.
historyPrevious states. A raise or reallocation updates the current fields and appends the previous state here with a dated note saying why. History is never deleted.

Each plan entry requires shared_ceiling_usd, authorized_by, recorded_utc. The plan names its packages and points here for their ceilings.

spend / LEDGER.md

Ledger

The ledger is append-only and includes failed and voided attempts: a run that billed before it failed still spent the money. Its table uses these columns, in order: Row | Date | Plan | Experiment | Attempt | Job | Phase | Committed USD | Actual USD | Outcome.

| Row | Date | Plan | Experiment | Attempt | Job | Phase | Committed USD | Actual USD | Outcome |
|---|---|---|---|---|---|---|---|---|---|
| S001 | 2026-09-10 | 2026-09-01-retrieval-comparison | EXP-002 | pilot-r1 | demo:SYN-PILOT-2 | pilot | 3.00 | 2.00 | complete |
ColumnMeaning
RowAn S-prefixed number, unique across the whole file.
PlanA plan filename stem, or - for spend that belongs to no plan, such as an owner-directed evaluation with no ceiling. A - row needs no authority entry.
Attempt, JobThe stable attempt identity the experiment record uses, and Backend:Job ID when the attempt ran on a runner.
Committed USDThe amount the authorization reserved.
Actual USDThe best current figure for what the attempt cost, or - while unknown.
OutcomeAn open or terminal outcome, described below.

An attempt holding its commitment has an open outcome: authorized or running. Afterward it takes a terminal outcome: complete, failed, canceled, operational-void, correction. Append /provisional to a terminal value when the actual is usage-estimated or not yet billed.

A wrong or late figure is corrected by a new correction row naming the same Attempt and carrying the difference, so the column sum stays the total. A final bill that adds 0.50 for idle time appends a 0.50 correction; it does not restate the attempt. A row with the wrong number of cells is an error, never skipped, because a skipped money row understates spend.

spend / deriving the account

Deriving the account

A, C, and U are computed when needed and never stored. For one plan at one as-of time, taking only rows whose Plan is that stem:

QuantityDerived from
A: incurredThe sum of Actual.
C: committedThe sum of Committed on rows whose Outcome is authorized or running.
U: plannedNot stored. The plan's remaining work in units, priced at a current rate when an authorization is built.

Headroom after commitments = shared ceiling − A − C, and it must still cover U. Neither headroom nor an underspend grants approval for more work. Unknown spend renders as unknown, never as zero, and an incomplete subtotal is not a settled total.

spend / run authorization

How a run gets its budget

Tooling that runs where the notebook is readable builds each authorization. It reads the authority and the ledger, checks the requested ceiling against headroom, converts the ceiling and the placement's rate into seconds (or calls, or tokens), writes the commitment as a ledger row, and hands the script an authorization in those units.

The script enforces the units it was given and refuses an authorization that carries any currency field, since a money field could carry a ceiling that has since changed. It never reads the notebook, which often does not reach the execution host, and holds no plan-level number of its own.

Cost estimation and tracking guide Plan resources and spend Experiment resources Reference