2026-08-29 · Updated 2026-08-29 · 7 min read

Repeated-attempt variance, observed and downloadable

Observed repeated-attempt variance from the public record: 18 declared coding-agent identities listed two or more times against one fixed 500-instance task set moved by a median 5.2 and up to 39.8 percentage points between listings — two charts, a hash-pinned downloadable JSON summary, and the disclosures that bound what these numbers mean.

By Juno AI INC · benchmark · evaluation · variance

Ask whether a coding agent is stable and the honest answer is a number: run it again and measure how far it moves. Public evidence for that number is scarce, because leaderboards rank submissions instead of re-listing one system repeatedly. This page publishes the observed version of the answer. Every declared coding-agent identity that the public SWE-bench Verified board's pinned 2026-08-29 snapshot records two or more times — eighteen identities holding forty-one listings against one fixed, human-validated task set of 500 real instances — appears below with the movement between its listings, in two charts, and as a versioned, hash-pinned JSON summary anyone can download and cite. Every figure on this page recomputes from the pinned parent extract, and the evidence date is 2026-08-29.

One boundary belongs up front. Elsewhere in this program, the variance guide works out what a small sample of attempts can certify, using model-based arithmetic, with figures computed from committed formulas. This page is deliberately not that. What follows are observed listings — what submitters published, when, under one declared name — and a re-listing is not a controlled re-run. Nothing here confirms, calibrates, or updates that guide's illustrative figures; the two treatments carry different kinds of truth and stay separate on purpose.

The unit: one declared identity, listed twice

The input is dataset version 1 of the longitudinal extract, fetched on 2026-08-29 from the board's public page (which reported itself last modified on 2026-08-10) and pinned by an entries hash, sha256:cbd2393a…8d02b6, so the derivation input is content-addressed exactly like the site's other evidence. The extract carries 180 dated entries. Grouping them by declared identity — the published agent and model display names, exactly as listed — yields 157 distinct identities, of which eighteen appear two or more times; those eighteen hold 41 of the 180 entries between them.

The unit matters more than the count. The board pins the task set, never the system: a submission is one evaluated run of whatever its submitter chose to run, identified only by the names the submitter typed. A repeated listing is therefore evidence of movement between submissions of one name — an upper bound on run-to-run variance, never a clean measurement of it. What else moves inside a name is recorded per group in the downloadable file and treated in the disclosures section below.

The spread, identity by identity

Observed resolved-percentage spread for 18 repeated declared identities on the SWE-bench Verified boardOne row per declared system identity submitted two or more times between 2024-05-09 and 2026-02-17. Each bar spans that identity's lowest to highest published resolved percentage of the 500-instance set; each dot is one listing. Spreads run from 0.8 to 39.8 percentage points with a median of 5.2. Rows are ordered by spread, largest first.range bar: an identity's lowest to highest listing · dot: one listing · label: spread in percentage points020406080100resolved percentage of the 500-instance setAmazon Q Developer Agent · Undisclosed39.8EPAM AI/Run Developer Agent · Claude 3.5 Sonnet23.2Gru · Undisclosed11.8Emergent E1 · Multiple10.6Nemotron-CORTEXA · Multiple10.0EntroPO + R2E · Qwen3-Coder-30B-A3B-Instruct8.2nFactorial · Multiple7.6Solver · Undisclosed6.4nFactorial · Undisclosed5.8Warp · Multiple4.6Refact.ai Agent · Multiple4.0mini-SWE-agent · GPT 5.23.8mini-SWE-agent · GPT 5 mini3.6EPAM AI/Run Developer Agent · GPT-4o3.0SWE-Fixer · Qwen2.5 (7B + 72B)2.6mini-SWE-agent · Claude 4.5 Opus2.4mini-SWE-agent · Claude 4.5 Sonnet0.8Salesforce AI Research SAGE · Multiple0.8
Each row is one declared system identity listed two or more times on the public board; bars span the lowest to highest published resolved percentage of that identity, dots mark individual listings. Derived from dataset version 1 of the longitudinal extract; evidence date 2026-08-29.

Every repeated identity moved: not one group published the same resolved percentage twice. Movement runs from four-fifths of a point to 39.8 points, with a median of 5.2. Nine of the eighteen groups stayed within 4.6 points of themselves; five moved ten points or more. The observed percentages inside these groups span 24.0 to 76.8.

The extremes teach the reading discipline. The smallest movement, 0.8 points, belongs to a pair of Salesforce AI Research SAGE listings whose folder identifiers on the board — …SAGE_bash_only against …SAGE_OpenHands — name different scaffolds for the two runs; the closest numbers among these eighteen groups are also not one system re-run. The largest, 39.8 points across four Amazon Q Developer Agent listings spanning 331 days, is exactly what a declared name permits: code, configuration, and provider models can all change under one display name while the task set stays fixed. Read the rows as ranges of published outcomes, and the range as a bound on stability — nothing narrower.

Spread against time between listings

Observed spread against days between the first and last listing of each repeated identityEach point is one of the 18 repeated declared identities; horizontal position is the span in days between its first and last listing date, vertical position its observed spread in percentage points. The largest spreads sit on the longest spans; the two hollow points mark groups holding two listings on one date.observed spread, percentage points010203040050100150200250300350days between the first and last listing39.88.23.8group with two listings on one datelistings on different dates
Each point is one repeated declared identity: horizontal position is the days between its first and last listing, vertical position its observed spread. Hollow points mark groups that hold two listings on one date. Evidence date 2026-08-29.

The second chart plots each identity's spread against the days between its first and last listing, and it carries the pattern honesty demands. The largest movements ride the longest spans — the top three spreads all sit on spans past three months — yet a long span guarantees nothing: one mini-SWE-agent pair waited 141 days between listings and moved 0.8 points. Two hollow points mark the groups that hold two listings on one date, and each sits where its group's span puts it: the EntroPO + R2E pair at the zero-day edge with 8.2 points, and the mini-SWE-agent with GPT 5.2 group at 68 days with 3.8 — its same-date entries are two of three listings, so the group's span reaches beyond them.

Those same-date pairs are the closest thing to a re-run the public surface offers, and neither is clean. One identity, mini-SWE-agent with GPT 5.2, carries three listings; two of them share the date 2025-12-11, at 69.0 and 71.8 — 2.8 points apart — with the same scaffold version (1.17.2), the same model identifier, and the same attempt policy. What differs, besides the outcome, is the reasoning-effort tag, high against none recorded, which each listing's folder identifier echoes, and the reported cost: $260.13 for the high-effort run against $134.83 for its partner, at 19.8 against 16.3 mean API calls per instance. The EntroPO + R2E pair shares one listing date, 2025-09-01, and one model, but its two listings carry different attempt policies (one against 2+), so 8.2 points separate two quantities that were not the same measurement to begin with. Even the best-matched pairs in the public record carry a residual difference the surface cannot explain away.

What a repeated listing cannot isolate

A declared name is not a pinned artifact, and the summary records per group exactly which identity-bearing fields changed inside it. The counts argue for humility:

  • Six of the eighteen groups publish differing model identifier arrays between listings. One Nemotron-CORTEXA listing names twelve model identifiers and its partner names nine; Refact.ai Agent went from three identifiers to two; Warp from four to two. A display name of Multiple can hold a recomposed ensemble.
  • All four mini-SWE-agent groups mix scaffold generations, 1.x with 2.0.0, and the dataset guide records the board's own caution that results across those generations are not necessarily comparable — so even the board's most controlled family of submissions changed scaffolds inside every repeated name.
  • Three groups mix attempt policies, so part of their spread is a changed measurement rather than a changed outcome.
  • Listing dates order publication, not measurement; every run happened before the listing that publishes it, by a lead the surface never states, and spans of days are spans of listing dates.
  • Every entry is third-party and self-reported; the verification marker travels per entry in the file, mixed, and nothing here upgrades an unchecked row.

Observed spread therefore bounds; it never isolates. Between two listings of one name sits whatever the submitters changed, whatever the providers changed underneath them, and ordinary sampling noise — and the public surface does not separate those. Anyone citing these numbers should cite them as what they are: movement between published listings of declared identities on one fixed task set, dated, hashed, and bounded.

Method, download, and citation

The summary is a pure derivation, and the code that derives it is committed beside the data. The rules: identity is the entry's published agent and model display names; a group requires at least two listings; spread is a group's highest minus its lowest published resolved percentage, computed in integer tenths so no floating-point artifact can reach a cited number; groups order by spread descending, a group's listings by date. The document carries the pinned parent hash, the per-group disclosures listed above, and its own content hash over the groups array, sha256:80500ffc…d1306, computed over the canonical key-sorted form.

The download is repeated-attempt-variance-v1.json. Citing means naming this page's canonical URL — shown in the panel that follows — together with the dataset version and the groups hash; a refresh of the parent extract publishes a new summary version with new hashes and never rewrites this one. The committed checker re-derives the file from the parent extract and fails on any drift:

bash
npm --prefix frontend run seo:variance:check

The groups hash also recomputes independently of this site's toolchain:

bash
jq -cjS '.groups' public/data/repeated-attempt-variance-v1.json | shasum -a 256

Where this evidence sits

Three owners share this ground. The statistics of repeated attempts — what a small sample can certify, and at what cost — belong to the variance guide, whose arithmetic is model-based by design; this page's numbers are observed listings, and neither kind of figure validates the other. The listing extract itself, all 180 entries with its own methodology and limitations, belongs to the longitudinal dataset page; everything here derives from its version 1 and inherits its caveats. Turning such evidence into a repeatable research practice — refresh cadence, versioning governance, citation policy for recurring assets — is a methodology treatment this program published on 2026-08-29 as the recurring-research methodology register. For repeated-attempt evidence under your own control, the product-side answer is YYLO Benchmark's retained-attempt machinery; the charts above are what the public record already holds, and the panel below is how to cite them.