YYLO Benchmark · stable 0.2.1

YYLO Benchmark documentation

Prepare reviewed cases, delegate independent attempts, and evaluate retained outputs with frozen checklists and separate result rows.

On this page

Published thin runner. The v3 record format is not a package version. Benchmark 0.2.0 introduced the breaking lifecycle; 0.2.1 adds checklists. Read the version changelog and benchmark study and evidence limits.

Install and verify

Node.js 20.10+, Git, tar and POSIX process groups. Trusted-host workspace hygiene, not a filesystem/account sandbox. Harness setup remains caller-owned.

Stable is the default. The latest published prerelease is an explicit choice, not a stable upgrade. Source version 0.2.1 is tracked separately from publication.

Stable 0.2.1
npm install --global '@yylo/benchmark@0.2.1'
Prerelease 0.1.1-rc.2 · opt in
npm install --global '@yylo/benchmark@0.1.1-rc.2'

Registry channels checked . A prerelease channel may be older than stable; compare the exact versions above.

Stable 0.2.1

Reviewed reusable cases

Prepare historical Ledger tasks or supplied prompts from an explicitly reviewed pre-solution base.

Ledger draft reads completed requirements without completion responses; it never guesses the historical base.

Fresh source snapshots exclude future Git history, controller metadata and declared answer paths.

Stable 0.2.1
yylo-benchmark case draft --ledger-task TASK_ID
Stable 0.2.1
yylo-benchmark case create --help

Boundary: Review is required; no adversarial host isolation or automatic answer detection is claimed.

Capability evidence

Published 0.2.1 npm archive README inspected; SHA-512 integrity and README SHA-256 retained. No live-provider test is claimed.

Verified exact releases: 0.2.1.

Package-owned source files: src/v2/workspace.ts src/v2/cli.ts README.md . Their fingerprints are checked by the frontend generator.

Stable 0.2.1

Delegated task and workflow attempts

Compare models, harnesses and configurations in independent fresh starting repositories.

YYLO Pi, command adapters and existing Workflow Runner are supported.

Selected-step comparisons execute an independent prefix through that step for each variant, then stop.

Workflows own sessions and dependencies. Errors are retained; no automatic compatibility repair, retry or resume.

Stable 0.2.1
yylo-benchmark run --help

Boundary: Prefix results measure earlier-step effects as well. Delegation grants no production authority.

Capability evidence

Published 0.2.1 npm archive README inspected; SHA-512 integrity and README SHA-256 retained. No live-provider test is claimed.

Verified exact releases: 0.2.1.

Package-owned source files: src/v2/adapters.ts src/v2/harness.ts README.md . Their fingerprints are checked by the frontend generator.

Stable 0.2.1

Independent later evaluations

New checks, judges and human assessments consume retained outputs without rerunning candidates.

No prebound evaluator catalog is required.

Evaluators receive copies of retained files; original attempts and previous evaluations stay unchanged.

Execution, checks, judge disagreements and evaluator errors are separate; no automatic combined verdict.

Stable 0.2.1
yylo-benchmark evaluate --help

Boundary: Malformed, oversized or unavailable evidence is an evaluation error, not a capability verdict.

Capability evidence

Published 0.2.1 npm archive README inspected; SHA-512 integrity and README SHA-256 retained. No live-provider test is claimed.

Verified exact releases: 0.2.1.

Package-owned source files: src/v2/evaluators.ts README.md . Their fingerprints are checked by the frontend generator.

Stable 0.2.1

Traceable comparison rows

Display partial results, individual assessments, cost coverage, integrity errors and disqualifications honestly.

JSON and Markdown tables identify the treatment, attempt, execution status and each evaluator.

Known answer exposure is recorded as disqualification without deleting prior evidence.

Legacy evidence is preserved but not automatically migrated or interpreted by the new runtime.

Stable 0.2.1
yylo-benchmark report --help
Stable 0.2.1
yylo-benchmark disqualify --help

Boundary: One-shot results are not reliability estimates. Reported cost is not billing; unknowns remain unknown.

Capability evidence

Published 0.2.1 npm archive README inspected; SHA-512 integrity and README SHA-256 retained. No live-provider test is claimed.

Verified exact releases: 0.2.1.

Package-owned source files: src/v2/cli.ts README.md . Their fingerprints are checked by the frontend generator.

Stable 0.2.1

Frozen criteria and deterministic loss

Evaluate retained outputs against explicit project and task checklists.

Case creation freezes normalized project/task criteria and their digest. Public criteria are visible to candidates; changing an input file cannot revise the case.

Every criterion needs pass, fail or unknown plus evidence. The runner computes equal-weight loss = failed / total. Any unknown produces null loss with insufficient_evidence; malformed assessments produce evaluation_error.

Later evaluation criteria replace the whole inherited checklist. Resupply both files to retain both. Reports mark criteria_changed; earlier attempts and evaluations remain untouched.

Stable 0.2.1
yylo-benchmark case create --help
yylo-benchmark evaluate --help
yylo-benchmark report --help

Boundary: Loss is partial quality, not production acceptance. Compare matching checklist/evaluator identities and keep execution, errors and disqualification separate.

Capability evidence

Published archive README and package metadata inspected; artifact integrity and README digest retained below. Documentation review is not a live-provider test.

Verified exact releases: 0.2.1.

Package-owned source files: README.md . Their fingerprints are checked by the frontend generator.

Methodology and historical evidence

The current lifecycle is case draft/create → run → evaluate → report, with append-only disqualify. A historical case needs review of original requirements, the pre-solution base and answer exclusions. New checks, judges and human assessments may evaluate retained outputs without candidate reruns or original-catalog prebinding. Execution, check failures, judge disagreement, evaluator errors and disqualification stay separate; there is no automatic winner or combined verdict.

Each workflow variant starts from the same initial input and runs its own prefix through the selected step, then stops. Results include upstream effects; this is not a measurement of that step alone or downstream continuation. The workflow/harness owns sessions, dependencies and errors; cross-harness compatibility is not guaranteed.

Trusted-host workspace/context hygiene is not enforced filesystem/account/network isolation. Retained input is not proof of literal provider-message delivery. Unknown cost remains unknown, and one-shot results are not reliability estimates.

Thin-runner command walkthrough · Fifteen-task implementation study. Old v1/v2 commands and evidence are historical, not migrated or rewritten by the redesign.

Troubleshooting

Start with the executable’s --version and --help. If a command is missing, compare its availability label with your installed version. Do not silently install a prerelease to make an example work.

Ledger and Benchmark can be used standalone. YYLO delegates enforce compatibility separately; newest packages are not automatically a compatible combination. For sandbox, credential or workspace failures, repair the named prerequisite before dispatch. Preserve logs and retained evidence; do not delete state to hide a failure.

Sources and release evidence

Capability review: 2026-10-07. Reviewed version: 0.2.1. Changes by version. Canonical package repository. Published availability below is verified for exact versions, never inferred from the current source version.

  • @yylo/benchmark 0.1.0 artifact
    Integrity and provenance

    sha512-PlaJQis5v9/wyCyXsQsIQyglAnj3Ceb8sGkCLOzxtJ/oGJHsq353+MRNGGtz54DbAJXT+5I3rDBktE0vn6dr4w==

    README SHA-256: fc167dfaeddd599a2d23de22ecb545f8b3094375fdeb840513cc710aa4f8eb00

  • @yylo/benchmark 0.1.1-rc.2 artifact
    Integrity and provenance

    sha512-XulCxkBbf9s+P9tdhFBGTYn0+H+IRMpliYLou9odNDuZgTiGGEeQ01RlUdl67CKiv9gWGviewR9vYYKyNSHbJw==

    README SHA-256: a34cc4c8ca4125577dde74b557deb3d121d7f4395c4cc287a003f54f13709005

  • @yylo/benchmark 0.2.1 artifact
    Integrity and provenance

    sha512-9zF6P8EZ6wwSemLsYSugC5ttnL4A2awO0ETDdoHECAOPgRIx1gmsp3/Ce0en6LzlMY7VN0ibOgM8kvpTZOMrfg==

    README SHA-256: 1e72b97ef16a2a7011d4089c1b0bd156d8f0d2e1b3bc407a57f97d5b08766429

Discover Skills for reusable agent procedures. Skills summarize intent here; the tagged repository owns the complete instructions.