Cost-aware planning and accounting

documentation/ guide and reference

Cost-aware planning chooses the cheapest instrument sufficient for a scientific decision. Accounting connects that choice to the resources actually used, including failed attempts and charges that arrive after the science is complete.

Keep the scientific design in the plan, attempt evidence in experiment records, and execution authority explicit. The execution guide distinguishes a read-only preview from permission to run work. Discussing a cheaper option grants neither execution nor permission to change the records.

Choose an instrument that can settle the decision

An instrument is the analysis, check, or experiment that supplies evidence for a decision. Start with the decision and the smallest observation that could change it. Compare reusing saved data, checking the measurement, running a pilot, collecting the full comparison, and obtaining independent confirmation. A cheaper run that cannot distinguish the alternatives is wasted spend.

Controls establish whether the measurement is interpretable. Valid uncertainty describes what the observations can support. Independent confirmation tests a frozen claim using evidence that was not used to select it. Cost reduction must preserve these requirements; dropping controls or reusing selection data as confirmation undermines the intended inference.

A fictional paired retrieval campaign

A team wants to decide whether a revised retrieval method improves answer quality over its baseline. A paired comparison evaluates both methods on the same questions. Existing outputs can support preliminary analysis, but they do not cover the frozen revision. Scoring checks and a pilot precede the full comparison. Independent confirmation uses held-out questions after the design is frozen.

Every identifier, date, usage record, and billing source below is synthetic. The campaign illustrates resource decisions, not scientific findings. No actual experiment records or provider quotations are represented.

CandidateDecision it can support
Existing-data analysisInspect saved paired outputs and uncertainty before collecting more. Preserve the selection history; reused evidence cannot independently confirm a claim selected from it.
Measurement checksUse known-answer cases to test the scorer. A failure stops interpretation and can avoid an expensive invalid comparison.
PilotCheck feasibility, resource use, and whether the instrument reaches the phenomenon. A small pilot does not replace the planned uncertainty analysis.
Full comparisonCollect the paired observations needed for the stated decision rule, with its controls and uncertainty requirements intact.
Independent confirmationTest the frozen claim on untouched evidence when confirmation is required. Reserve the evidence and resources before exploratory choices consume them.

Compare total costs, including data preparation, analysis, and human review. Existing-data analysis can still use compute and labor. A failed check may make stopping the cheapest sufficient action. Record what the cheaper option cannot answer before choosing a larger run.

Estimate setup, attempts, and conditional work

A baseline is the dated estimate for the intended campaign under stated assumptions. An attempt is one execution, including an execution that fails. A contingency is capacity inside the approved ceiling for specified uncertainty or failure; it is not already spent. Estimate setup, successful attempts, failures and retries, teardown, and any conditional branches separately.

State the quantity, unit, rate, currency, rate date, and billing scope. In rental accounting, one GPU provisioned for one hour consumes one GPU-hour; two GPUs consume two GPU-hours. Compute quantity, elapsed time, money, and human effort answer different questions. Parallel execution can shorten elapsed time without reducing GPU-hours. Queue waits can lengthen elapsed time without being billable compute. Human review time is neither of those quantities.

The campaign baseline

The synthetic estimate is dated 2026-09-09. Each attempt uses one GPU at a flat, invented rate of USD 2 per provisioned GPU-hour, including provisioned idle time. This is a teaching rate, not a provider quote. The scope is GPU rental only. CPU work outside that rental, storage, network transfer, taxes, and human labor are unpriced and outside this ceiling; their cost is unknown, not zero. An all-in budget would need those items priced and approved separately.

Planned workGPU-hoursUSD
Setup and scoring checks0.51.00
Pilot12.00
Comparison: four attempts4 × 1 = 44 × 2 = 8.00
Independent confirmation36.00
Baseline0.5 + 1 + 4 + 3 = 8.517.00
Contingency24.00
Proposed ceiling, including contingency8.5 + 2 = 10.517.00 + 4.00 = 21.00

The baseline assumes the pilot succeeds on its first attempt and the scientific gates permit comparison and confirmation. Confirmation is planned but conditional on those gates. The full-path estimate includes it. A stop after a failed gate is a different branch with its own cost; no outcome probability has been assigned to either branch.

The 2 GPU-hour contingency allows the specified retry and provisioned-idle uncertainty. It does not add new comparison conditions. The 3 GPU-hours for independent confirmation are already in the baseline and fund the test. The held-out questions form the scientific confirmation reserve; monetary contingency cannot substitute for untouched evidence.

Estimate elapsed time separately using dependencies, concurrency, queue delays, and review waits. Estimate human effort for setup, diagnosis, analysis, and review in person-hours. Neither is estimated numerically in this GPU-rental example, so 8.5 GPU-hours is not a completion-time or labor promise. For real rates, record billing granularity, minimum charges, startup and shutdown rules, and whether discounts or credits apply.

When attempt duration or price is uncertain, retain a range or a conservative bound and explain its basis. A pilot measurement is evidence for a revision, not a guarantee that every full run will cost the same. List conditional paths without inventing probabilities to make a precise-looking expected total.

Preview costs without changing records

Replace PLAN_PATH with the plan path and send this prompt to the coding agent. The preview may discuss already recorded costs and calculate a proposed rollup in its reply. It cannot reconcile the notebook or launch work.

Use research-lab-notebook to preview the costs of PLAN_PATH using existing records and saved outputs only. Compare existing-data analysis, measurement checks, a pilot, the full comparison, and required independent confirmation against the scientific decision. Preserve controls and valid uncertainty. Report setup, attempt counts, failures and retries, conditional branches, rates, currency, rate dates, billing scope, and uncertainty. Keep money, compute, elapsed time, and human effort separate. Identify incurred usage, remaining commitments, planned uncommitted work, contingency, approved limits, and missing evidence as of a stated time. Show the arithmetic without treating unknowns as zero. Do not modify files, reconcile records, launch or retry jobs, or change execution authority.

Authorize a scope and a ceiling together

An approved ceiling limits a specified resource within a stated scope. Approval names the work, units, limits, permitted retries, scientific gates, and who may decide to proceed. Preserve the dated baseline and dated approval rather than replacing the baseline whenever estimates change. A larger ceiling or cheaper forecast does not authorize a different study.

Keep this in the plan's optional human-readable Budget and accounting section: estimates and assumptions, approved scope and limits, contingency, retry policy, dated approvals, and links to rollups with an explicit as-of time. The heading is a convention for readers, not a required schema field or a validator-enforced spending control.

The campaign's bounded approval

A synthetic approval on 2026-09-09 covers setup, the pilot, four comparison attempts, and independent confirmation after their scientific gates. It sets two ceilings: 10.5 GPU-hours and USD 21, each including contingency. A single pilot retry is permitted after diagnosing and recording an execution failure, with at most 1 additional GPU-hour. Another retry, a changed design, or work beyond either ceiling requires a new decision. These statements illustrate approval; they authorize no real job.

Before submitting parallel jobs, reserve conservative upper bounds for their remaining charges together. Include permitted retries, provisioning overhead, billing increments, and the delay before a stop takes effect. Several agents cannot each spend the same apparent balance. Coordinate reservations against one current plan account and record who controls admission. A central reservation here is an operating responsibility, not a new notebook ledger or a claimed scheduler feature.

A written limit does not stop a process. Use the runner's actually supported duration, concurrency, or spending controls where available, and describe their limitations in RUNNER.md. When hard enforcement is unavailable, use conservative admissions and explicit checkpoints rather than promising a hard cap. Continuous execution still respects the approved scope, resource limits, and scientific gates.

Track incurred cost and remaining work separately

Experiment records own the evidence for each attempt. Put usage and cost in the optional human-readable Cost and resources section beside Runs. Use the existing Backend plus Job ID to identify each attempt. A manual run needs a stable attempt identity too. Retrying creates another attempt whose cost remains visible even if it produces no usable scientific output. No new Runs columns are required.

Record the quantity and unit, cost and currency, source identity, calculation basis, coverage period, and retrieval timestamp. Provider-reported cost comes from a provider's usage or billing source. Usage-estimated cost is inferred from usage and a stated rate. Either can be provisional. Mark whether a reported value is billed or provisional separately from how it was obtained; a usage dashboard is not necessarily a settled invoice.

RUNNER.md describes the supported usage and billing CLI or API sources, their units, coverage, update delays, and limitations. Use those interfaces and retained evidence, not tool-private databases. If a source supplies runtime but no charge, label the calculated cost as estimated. If usage or billing is missing, retain the gap and a responsible follow-up owner. Unknown cost is not a zero-cost run.

Use one as-of time and one accounting scope

For each unit and scope, split the forecast into three amounts. A is incurred cost or usage to date. C is the additional forecast for work already running or submitted, excluding its incurred portion. U is the forecast for planned work not yet committed. The forecast is A + C + U. State the as-of time and uncertainty of each component. Use separate totals for currencies or scopes unless an explicit conversion and attribution basis supports combining them.

The campaign at its checkpoint

At synthetic time 2026-09-10T12:00Z, setup is complete. The first pilot failed after 0.25 GPU-hour; its authorized retry completed in 1 GPU-hour. Two comparison attempts finished. Two more have each consumed 0.25 GPU-hour and are expected to need 0.75 more. Confirmation has not been submitted.

Every attempt below uses synthetic backend demo; the row labels are its synthetic Job IDs. The saved runtime snapshot SYN-USAGE-01, retrieved at that checkpoint, covers observed execution time through the same instant. Costs are usage-estimated at USD 2 per GPU-hour and provisional, not provider-reported billing. Provisioned idle time is not yet available from this source.

AttemptIncurred A: GPU-hours / USDAdditional committed C: GPU-hours / USD
SYN-SETUP0.5 / 1.000 / 0.00
SYN-PILOT-1, failed0.25 / 0.500 / 0.00
SYN-PILOT-2, retry1 / 2.000 / 0.00
SYN-COMP-11 / 2.000 / 0.00
SYN-COMP-21 / 2.000 / 0.00
SYN-COMP-3, running0.25 / 0.500.75 / 1.50
SYN-COMP-4, running0.25 / 0.500.75 / 1.50
Known subtotal4.25 / 8.501.5 / 3.00

Incurred usage is 0.5 + 0.25 + 1 + 1 + 1 + 0.25 + 0.25 = 4.25 GPU-hours. The running attempts need 0.75 + 0.75 = 1.5 additional GPU-hours. Planned, unsubmitted confirmation is U = 3 GPU-hours / USD 6. It has no Job ID yet and is not a commitment.

The forecast from known usage is 4.25 + 1.5 + 3 = 8.75 GPU-hours, or USD 8.50 + USD 3 + USD 6 = USD 17.50. On that incomplete basis, headroom after commitments is limit − A − C: 10.5 − 4.25 − 1.5 = 4.75 GPU-hours, or USD 21 − USD 8.50 − USD 3 = USD 9.50. It still has to cover the unsubmitted confirmation's 3 GPU-hours / USD 6.

The apparent forecast margin is 10.5 − 8.75 = 1.75 GPU-hours, or USD 21 − USD 17.50 = USD 3.50. The failed pilot has used 0.25 GPU-hour / USD 0.50 beyond baseline. Adding the original USD 4 contingency to the forecast again would double count part of that allowance. These figures exclude the disclosed, still unknown idle usage; admission needs a conservative allowance for that exposure.

As a running attempt consumes resources, move that portion from remaining C into incurred A; do not add its entire projected lifetime cost to A. When approved planned work is submitted, move its remaining forecast from U to C. A cancellation request reduces C only after cancellation is confirmed. Keep cancellation fees, consumption before shutdown, and unknown billing visible. Free headroom and underspend never grant scope approval.

Shared setup, storage, or provider charges need an explicit attribution policy and source identity. Allocate a charge once, or leave it visibly unallocated. Link multiple experiments to that evidence without charging each for its full value. A plan rollup links the owning experiment accounts at a stated as-of time; it does not turn copied totals into new spend.

Reconcile final billing without reopening execution

Reconciliation matches later usage and billing evidence to the attempt accounts and resolves discrepancies. Financial settlement means the scoped account has its final charges and resolved allocations. Scientific processing may finish earlier, provided available cost evidence, missing charges, and a follow-up owner are recorded. Mark financial settlement pending until those gaps are resolved.

The campaign's provisional close and final settlement

The two running comparisons each finish using their remaining 0.75 GPU-hour. After the required gate and within the existing synthetic approval, confirmation is submitted as demo + SYN-CONFIRM and completes in 3 GPU-hours. Those quantities transfer from C and U into A. No scientific result is asserted here.

Runtime snapshot SYN-USAGE-02, retrieved as of 2026-09-11T18:00Z, gives the known portion of A = 4.25 + 1.5 + 3 = 8.75 GPU-hours / USD 17.50, with C = U = 0. That usage-estimated close is provisional. The campaign owner retains responsibility for matching the final GPU-rental bill; billing for provisioned idle time remains pending.

Synthetic final bill SYN-BILL-01, retrieved on 2026-09-12T12:00Z and covering the campaign's full rental interval, reports 9 GPU-hours / USD 18 billed. Its line for demo + SYN-SETUP includes 0.25 GPU-hour of provisioned idle time missing from the runtime snapshots. Setup is corrected from 0.5 GPU-hour / USD 1 to 0.75 GPU-hour / USD 1.50. The remaining attempt charges agree.

Account versionGPU-hoursUSD
Original baseline8.517.00
Provisional close, usage-estimated8.7517.50
Final close, provider-reported and billed918.00
Final variance from baseline+0.5+1.00
Unused approved ceiling10.5 − 9 = 1.521.00 − 18.00 = 3.00

Final A is 0.75 + 0.25 + 1 + 4 + 3 = 9 GPU-hours; C = U = 0. The USD 1 baseline variance is the failed pilot's USD 0.50 plus the previously missing idle charge's USD 0.50. Retain the USD 17.50 estimate with its date and mark it superseded by USD 18. The final bill replaces the estimate; it is not USD 18 of additional spend.

Financial settlement is complete for this synthetic GPU-rental scope. The unpriced categories outside it remain outside the claim. The unused USD 3 authorizes no further runs. A successor estimate should include provisioned idle overhead and the observed retry exposure, without treating one failure as an estimated failure probability.

Correct evidence at its owning experiment account, retain the old value and reason for supersession, then update the linked plan rollup. Compare the settled total with the original baseline, explain variance, and state outstanding charges or allocations. If a provider later revises a bill, record the new evidence and reopen financial settlement as needed; scientific closure does not make the earlier charge immutable.

Authorize an accounting-only update

Use this separate request when you want the records changed. Replace PLAN_PATH, EXPERIMENT_PATHS, SOURCE_PATHS, and AS_OF with the plan, owning experiment files, saved usage or billing evidence, and cutoff timestamp. This authorizes accounting edits, not experiment execution or a larger budget.

Use research-lab-notebook to update only the accounting in PLAN_PATH and EXPERIMENT_PATHS, using existing evidence in SOURCE_PATHS as of AS_OF. Match each charge to Backend plus Job ID or its stable manual-attempt identity. Record usage, cost, currency, source identity, basis, coverage, and retrieval timestamp. Distinguish provider-reported from usage-estimated amounts and billed from provisional amounts. Preserve prior estimates as superseded, not additional spend. Apply the recorded shared-cost attribution policy once and disclose unallocated or unknown charges. Update the linked plan rollup with incurred A, additional committed C, planned uncommitted U, A + C + U, and headroom against the existing approved ceiling, including its contingency. Explain baseline variance and pending settlement; name the follow-up owner or flag that one is missing. Preserve scientific records and existing scope, approvals, and limits. Do not launch, submit, retry, cancel, or rerun jobs, collect new scientific evidence, approve new work, or treat this update as execution approval.

A read-only results walkthrough may explain these costs from existing records. Writing corrections requires separate authorization such as the accounting request above. Any proposed follow-up returns to the plan workflow for a scientific and resource decision.