2026-08-29 · Updated 2026-08-29 · 7 min read
Repeated-attempt variance, observed and downloadable
Observed repeated-attempt variance from the public record: 18 declared coding-agent identities listed two or more times against one fixed 500-instance task set moved by a median 5.2 and up to 39.8 percentage points between listings — two charts, a hash-pinned downloadable JSON summary, and the disclosures that bound what these numbers mean.
By Juno AI INC · benchmark · evaluation · variance
Ask whether a coding agent is stable and the honest answer is a number: run it again and measure how far it moves. Public evidence for that number is scarce, because leaderboards rank submissions instead of re-listing one system repeatedly. This page publishes the observed version of the answer. Every declared coding-agent identity that the public SWE-bench Verified board's pinned 2026-08-29 snapshot records two or more times — eighteen identities holding forty-one listings against one fixed, human-validated task set of 500 real instances — appears below with the movement between its listings, in two charts, and as a versioned, hash-pinned JSON summary anyone can download and cite. Every figure on this page recomputes from the pinned parent extract, and the evidence date is 2026-08-29.
One boundary belongs up front. Elsewhere in this program, the variance guide works out what a small sample of attempts can certify, using model-based arithmetic, with figures computed from committed formulas. This page is deliberately not that. What follows are observed listings — what submitters published, when, under one declared name — and a re-listing is not a controlled re-run. Nothing here confirms, calibrates, or updates that guide's illustrative figures; the two treatments carry different kinds of truth and stay separate on purpose.
The unit: one declared identity, listed twice
The input is dataset version 1 of the longitudinal extract, fetched on 2026-08-29 from the board's public page (which reported itself last modified on 2026-08-10) and pinned by an entries hash, sha256:cbd2393a…8d02b6, so the derivation input is content-addressed exactly like the site's other evidence. The extract carries 180 dated entries. Grouping them by declared identity — the published agent and model display names, exactly as listed — yields 157 distinct identities, of which eighteen appear two or more times; those eighteen hold 41 of the 180 entries between them.
The unit matters more than the count. The board pins the task set, never the system: a submission is one evaluated run of whatever its submitter chose to run, identified only by the names the submitter typed. A repeated listing is therefore evidence of movement between submissions of one name — an upper bound on run-to-run variance, never a clean measurement of it. What else moves inside a name is recorded per group in the downloadable file and treated in the disclosures section below.
The spread, identity by identity
Every repeated identity moved: not one group published the same resolved percentage twice. Movement runs from four-fifths of a point to 39.8 points, with a median of 5.2. Nine of the eighteen groups stayed within 4.6 points of themselves; five moved ten points or more. The observed percentages inside these groups span 24.0 to 76.8.
The extremes teach the reading discipline. The smallest movement, 0.8 points, belongs to a pair of Salesforce AI Research SAGE listings whose folder identifiers on the board — …SAGE_bash_only against …SAGE_OpenHands — name different scaffolds for the two runs; the closest numbers among these eighteen groups are also not one system re-run. The largest, 39.8 points across four Amazon Q Developer Agent listings spanning 331 days, is exactly what a declared name permits: code, configuration, and provider models can all change under one display name while the task set stays fixed. Read the rows as ranges of published outcomes, and the range as a bound on stability — nothing narrower.
Spread against time between listings
The second chart plots each identity's spread against the days between its first and last listing, and it carries the pattern honesty demands. The largest movements ride the longest spans — the top three spreads all sit on spans past three months — yet a long span guarantees nothing: one mini-SWE-agent pair waited 141 days between listings and moved 0.8 points. Two hollow points mark the groups that hold two listings on one date, and each sits where its group's span puts it: the EntroPO + R2E pair at the zero-day edge with 8.2 points, and the mini-SWE-agent with GPT 5.2 group at 68 days with 3.8 — its same-date entries are two of three listings, so the group's span reaches beyond them.
Those same-date pairs are the closest thing to a re-run the public surface offers, and neither is clean. One identity, mini-SWE-agent with GPT 5.2, carries three listings; two of them share the date 2025-12-11, at 69.0 and 71.8 — 2.8 points apart — with the same scaffold version (1.17.2), the same model identifier, and the same attempt policy. What differs, besides the outcome, is the reasoning-effort tag, high against none recorded, which each listing's folder identifier echoes, and the reported cost: $260.13 for the high-effort run against $134.83 for its partner, at 19.8 against 16.3 mean API calls per instance. The EntroPO + R2E pair shares one listing date, 2025-09-01, and one model, but its two listings carry different attempt policies (one against 2+), so 8.2 points separate two quantities that were not the same measurement to begin with. Even the best-matched pairs in the public record carry a residual difference the surface cannot explain away.
What a repeated listing cannot isolate
A declared name is not a pinned artifact, and the summary records per group exactly which identity-bearing fields changed inside it. The counts argue for humility:
- Six of the eighteen groups publish differing model identifier arrays between listings. One Nemotron-CORTEXA listing names twelve model identifiers and its partner names nine; Refact.ai Agent went from three identifiers to two; Warp from four to two. A display name of Multiple can hold a recomposed ensemble.
- All four mini-SWE-agent groups mix scaffold generations, 1.x with 2.0.0, and the dataset guide records the board's own caution that results across those generations are not necessarily comparable — so even the board's most controlled family of submissions changed scaffolds inside every repeated name.
- Three groups mix attempt policies, so part of their spread is a changed measurement rather than a changed outcome.
- Listing dates order publication, not measurement; every run happened before the listing that publishes it, by a lead the surface never states, and spans of days are spans of listing dates.
- Every entry is third-party and self-reported; the verification marker travels per entry in the file, mixed, and nothing here upgrades an unchecked row.
Observed spread therefore bounds; it never isolates. Between two listings of one name sits whatever the submitters changed, whatever the providers changed underneath them, and ordinary sampling noise — and the public surface does not separate those. Anyone citing these numbers should cite them as what they are: movement between published listings of declared identities on one fixed task set, dated, hashed, and bounded.
Method, download, and citation
The summary is a pure derivation, and the code that derives it is committed beside the data. The rules: identity is the entry's published agent and model display names; a group requires at least two listings; spread is a group's highest minus its lowest published resolved percentage, computed in integer tenths so no floating-point artifact can reach a cited number; groups order by spread descending, a group's listings by date. The document carries the pinned parent hash, the per-group disclosures listed above, and its own content hash over the groups array, sha256:80500ffc…d1306, computed over the canonical key-sorted form.
The download is repeated-attempt-variance-v1.json. Citing means naming this page's canonical URL — shown in the panel that follows — together with the dataset version and the groups hash; a refresh of the parent extract publishes a new summary version with new hashes and never rewrites this one. The committed checker re-derives the file from the parent extract and fails on any drift:
The groups hash also recomputes independently of this site's toolchain:
Where this evidence sits
Three owners share this ground. The statistics of repeated attempts — what a small sample can certify, and at what cost — belong to the variance guide, whose arithmetic is model-based by design; this page's numbers are observed listings, and neither kind of figure validates the other. The listing extract itself, all 180 entries with its own methodology and limitations, belongs to the longitudinal dataset page; everything here derives from its version 1 and inherits its caveats. Turning such evidence into a repeatable research practice — refresh cadence, versioning governance, citation policy for recurring assets — is a methodology treatment this program published on 2026-08-29 as the recurring-research methodology register. For repeated-attempt evidence under your own control, the product-side answer is YYLO Benchmark's retained-attempt machinery; the charts above are what the public record already holds, and the panel below is how to cite them.