This page is generated from docs/proof/claims.toml
and the evidence files it points at. It is not written by hand. The same
generator runs in the project's health check, so a number that drifts from its
evidence fails the build rather than reaching a reader.
| Claim | Value | Evidence |
|---|---|---|
| Machine time for one deployed run of the whole lifecycle machine time only; the human approval is timed separately and excluded | 55.3 s of machine time | spikes/judge_run/evidence.json |
| The single human approval the same run required | 47.7 s human approval | spikes/judge_run/evidence.json |
| The budget one deployed run must finish inside | 130s | spikes/judge_run/evidence.json |
| The same work done by hand, timed author-timed, not practitioner-reviewed | 663.5 s | spikes/manual_baseline/evidence.json |
| Steps in the hand-done walkthrough author-timed, not practitioner-reviewed | 20 steps | spikes/manual_baseline/evidence.json |
| Steps of that walkthrough the run removes measured against an author-timed baseline, not a practitioner-reviewed one | 19 of which the run removes | spikes/judge_run/evidence.json |
| Business days the two cases span the lifecycle clock is compressed, and the compression is disclosed on screen | 380 simulated business days | spikes/judge_run/evidence.json |
| Contracts re-verified at capture time | nine contracts | spikes/core_contracts/evidence.json |
| Graded domain cases, all passing a deterministic pass metric; re-run the suite before citing it | 24/24 | spikes/domain_evals/evidence.json |
| Mean score across the graded domain suite a deterministic pass metric; re-run the suite before citing it | 100% | spikes/domain_evals/evidence.json |
| Fields entered into the ERP without a human retyping them | 22 fields | spikes/judge_run/evidence.json |
| Days a non-compliant supplier was actually held from purchasing claimed only where both the hold and its release executed in the ERP | 5 enforced hold days | spikes/judge_run/evidence.json |
| Decisions policy requires a human to make | 1 policy-required intervention | spikes/judge_run/evidence.json |
| Duplicate ERP writes after a retry | 0 duplicate writes after a retry | spikes/judge_run/evidence.json |
| Contract and unit tests re-executed by the run that reports them re-executed by the harness, not quoted from a previous run | 549 passed | spikes/judge_run/evidence.json |
| Gross running cost for the whole project, month to date before credits, measured from the billing account rather than from published rates | $16.63 | spikes/cost_posture/evidence.json |
| Cloud credit remaining console-only; no API exposes the balance, so it is passed in when the measurement is taken | $138.09 | spikes/cost_posture/evidence.json |
| Measured daily uptime of the screening VM | 5.6 h/day | spikes/cost_posture/evidence.json |
| The ERP hold was written minutes before the rejection committed | read the evidence file | spikes/hitl_reject/evidence.json |
| The graded suite passed 8/8 when the generation pin was measured | read the evidence file | spikes/domain_evals/evidence.json |
| Reasoning tokens fell to effectively none and the timed sequence shortened | read the evidence file | spikes/gemini_37_eval/evidence.json |
| The build piece states the graded suite passed 8/8 with reasoning pinned off | read the evidence file | spikes/domain_evals/evidence.json |
| The build piece states a two-supplier run dropped from 85 to 57 seconds | read the evidence file | spikes/gemini_37_eval/evidence.json |