Source changelog · September 24, 2026
Benchmark changelog: a thin experiment runner
A smaller experiment loop, with explicit evidence limits rather than a combined verdict.
What changed
The loop is now case draft/create → run → evaluate → report, with append-only disqualify. Review historical Ledger requirements and the full pre-solution range once, or supply your own prompt. Each variant starts from an answer-free fresh Git snapshot, without future solution history in its prepared input.
Treatments select a model, harness and configuration. YYLO Pi, generic command adapters and the existing Workflow Runner are supported. Workflow/harness setup owns dependencies, session reuse and errors; Benchmark does not repair incompatible workflows or translate opaque session state.
For prefix comparisons, each variant starts from the same initial input and runs independently through the selected step, then stops. Results include upstream effects: they do not isolate that step's contribution and do not imply downstream continuation.
New checks, judges and human assessments can evaluate retained outputs later without candidate reruns or original-catalog prebinding. Execution status, deterministic checks, judge disagreement, evaluator errors, unknown costs and disqualification remain separate. There is no automatic combined verdict, ranking or winner.
Breaking transition — historical APIs
The old v1/v2 writable APIs and configuration, plan/recover/doctor/regrade/rejudge commands, plugin registry and governed-production boundary are retired. Ordinary workflow support is not removed: execution delegates to the existing runner. Delegation grants no production authority.
Do not point old plans at the new runtime. CLI 0.2.10 source requires Benchmark 0.2.0 exactly and Skills ^2.1.0; older Benchmark executables are rejected before dispatch. Updating source or installing the CLI does not replace installed skills. Preserve old evidence and its pinned implementation; it is not automatically migrated, deleted or reinterpreted. Create reviewed cases and explicit new output directories for new experiments. An unfinished intent is reported as interrupted_or_running, not automatically retried or resumed.
Evidence and isolation limits
This is trusted-host workspace/context hygiene, not enforced filesystem/account/network isolation. Candidate processes inherit host authentication and may access other host paths or public answers. Review answer exclusions; no general cheating prevention is claimed. Known answer exposure disqualifies an attempt without deleting evidence.
Retained input proves supplied harness input, not literal provider-message delivery. Prompt preprocessing may change it. Missing cost is unknown, not zero. A one-shot outcome is not a reliability estimate, and varying harnesses or configuration prevents model-only attribution.
The implementation reduced TypeScript source from 9,534 to 724 lines (92.4%). That measures source lines, not speed, quality, or model performance. Synthetic tests and packed local acceptance did not constitute a live model study.
Sources and next steps
The merged source and diff, pinned README, and pinned capability manifest own these claims.
The links above preserve the original redesign baseline. Current source version and compatibility facts are generated from the package manifests and pinned Skills source; they do not establish registry publication.
Use the current source walkthrough and version-labeled documentation. The ten-task retrospective remains historical evidence, not results for this rewrite. Publication, global installation and deployment are separate maintainer actions.