Execution · Source only

benchmark-yylo

Plan and run YYLO Benchmark studies (yylo-benchmark), reconstruct historical Ledger tasks, compare coding agents, and retain honest per-task evidence with explicit pilot approval gates.

Review before installation

This source skill requires a separately reviewed compatible release before installation. It is not included in the current published Skills version or YYLO’s seven-skill installer. Source merge does not install, activate, or publish it.

Its procedure gates model studies on exact-shape setup canaries, pilot approval, canonical dispatch, and immutable per-task evidence. Running candidates or judges requires separate authorization.

Review canonical source →

← Return to all Skills