Evaluation evidence / longitudinal comparison · v0.1.0-rc.7

Compare agent attempts with evidence that lasts.

YYLO Benchmark validates immutable plans, runs isolated attempts, governs grading, reconciles evidence, and produces bounded reports without hiding source artifacts.

Supplementary visual overview

YYLO Benchmark at a glance.

Product behavior and commands are defined by the verified HTML documentation; this artwork provides a high-level visual summary.

Open full-size image Visual overview of the YYLO Benchmark planning, execution, grading, reporting, and evidence pipeline

Install / first run

Start with one bounded outcome.

Install the synchronized release, then begin with a change small enough to validate and review.

Install
npm install -g @yylo/benchmark@0.1.0-rc.7
First run
yy benchmark --help

Available today / mechanisms

Control is visible, not implied.

01

Immutable plans

Bind cases, models, attempts, source identities, and constraints before execution.

02

Isolated attempts

Keep candidate runs independent and retain execution and cost provenance when available.

03

Governed grading

Use pinned graders and blinded evaluation rather than untraceable model preference.

04

Recovery and reports

Reconcile durable intent, recover safely, and compare retained evidence over time.

Product direction / clearly separated

How this layer strengthens the System of Work.

Benchmark evidence is designed to help teams learn which agent and workflow works best for a repository and change type. Automatic production workflow selection requires connected real-world outcome history and is not a current-release claim.

Next step

Use this layer inside YYLO.

Explore the YYLO control plane