2026-08-29 · Updated 2026-08-29 · 9 min read
The state of coding-agent harnesses, edition 1
A recurring, evidence-dated report on the coding-agent harness landscape: five scaffold-layer findings recomputed from a pinned 180-entry public leaderboard extract, two product-layer observations drawn from an evidence-dated matrix, and the methodology, changelog, and update ownership that keep it maintainable.
By Juno AI INC · agent-harness · landscape-report
This is edition 1 of a recurring, evidence-dated report on the coding-agent landscape — both the scaffolds that run agents on public evaluation boards and the harness products that host them for developers. It was assembled on 2026-08-29 from two instruments whose provenance this site controls: a pinned, hash-identified extract of a public coding-agent leaderboard, and an evidence-dated harness matrix published on this site. The short version: the environment around coding agents is now a multi-product category spanning IDEs, terminals, and control planes; underneath it, the public-submission record widened through 2025 — 55 distinct scaffold names appear in that year's stream — and then standardized, with every entry dated 2026 running one fixed evaluation scaffold; recurring scaffolds routinely wear more than one vendor's models; and cost disclosure, absent from the 2024 stream, rides along with every 2026 entry while verification coverage still lags. Every scaffold-layer number below recomputes from the pinned extract. Nothing on this page is a search-demand estimate, a market-share figure, or an adoption measurement.
The reader this report serves is an engineering leader tracking the harness landscape — someone who needs dated, checkable observations rather than vendor narratives, and who will be citing or reusing them. That is why the report ships as a numbered edition with a stated methodology, a changelog, and a declared update owner: a landscape claim is only worth citing if you can name the bytes it came from and the date they were read.
What this report is — and is not
This report borrows its vocabulary from the five-layer boundary this site's harness-definition guide defends: model, agent, harness, IDE, and control plane. Held straight, that boundary splits today's usage into two layers this report covers separately. Leaderboard scaffolds — SWE-agent, OpenHands, the board's own mini-SWE-agent — sit at the agent layer: they are the loops and tool surfaces that get evaluated on public task sets. Harness products — the CLIs, IDEs, and control planes developers actually run — sit at and above the harness layer. Casual usage stretches one word across both; a report that wants to stay maintainable cannot, so each layer gets its own instrument and its own findings below.
Three things this report deliberately is not. It is not a ranking: which harness handles which job better belongs to comparison pages, and this report states the shape of the layer, never a winner. It is not a support ledger: which harness currently talks to which agent is a facts page's job, and the versioned support and interoperability matrix has held that ground since publishing on this report's own evidence date — current-support questions wait for that page rather than this one. And it is not a market measurement: no download count, share estimate, or demand figure enters any finding, because this route's approved record carries no such coverage and none was invented.
Method: two instruments and an admissibility rule
The scaffold-layer instrument is a committed extract of the SWE-bench Verified board's public leaderboard, published on this site as the longitudinal leaderboard dataset: dataset version 1, holding 180 dated entries between 2023-10-10 and 2026-02-26 — every entry one evaluated agent system run against one fixed, human-filtered population of 500 instances drawn from real GitHub issues — kept in published order and sealed with a content hash over the entries (sha256:cbd2393a235c0276211fb4d7035c943ef7eb7e4be4902be99a5733f6380d02b6). The surface behind it was fetched on 2026-08-29 and carried a server last-modified stamp of 2026-08-10. This report reads the pinned extract rather than the live board, so a citation names the exact bytes behind a finding; the extract's field notes and limitations travel with the dataset page.
The product-layer instrument is the harness comparison matrix published on this site, dated 2026-08-28, covering seven shipping harnesses: Cursor, Kiro, OpenCode, Claude Code, Codex, Pi, and YYLO. Product-layer observations below must be checkable against that matrix's rows at its date; its verdicts are restated nowhere on this page.
The admissibility rule is mechanical. A numeric finding enters only if it recomputes from the pinned extract at the version named above; a product-layer observation enters only if it is checkable against the matrix at its evidence date; anything requiring a fresh external fetch beyond those two instruments waits for a future edition. The inherited limits are stated once here and bind every finding: every entry is a third party reporting its own run on one fixed task set, so a missing system means no one submitted it, not that it cannot do the work; an entry's date is the date it was listed, and the run behind it happened earlier by an unknown margin; and the population is one board's Verified task set, not a census of practice. The recipe that reproduces the population counts:
The remaining scaffold-layer findings read the same file's verified, agent_version, open_source_system, agent_org, and model_org fields; no scaffold-layer number comes from anywhere else.
The scaffold layer at this evidence date
- The name layer is wide and mostly single-use. The stream publishes 180 entries wearing 77 distinct agent/scaffold names, submitted under 40 distinct organizations as published (51 entries carry no organization). Most names are one-offs: 52 appear in exactly one submission, 25 recur. Two-thirds of the stream — 116 of 180 entries — carries the board's open-source-system marker.
- Diversity rose through 2025, then the stream standardized. One name covers the 2023 stream (4 entries); 2024 publishes 51 entries under 29 names; 2025 publishes 112 under 55; and 2026 publishes 13 entries under a single name — the board's own fixed scaffold, which every one of those 13 entries carries as its frozen-scaffold marker. The 2026 shape is not convergence on a winning product; it is standardization of the evaluation scaffold itself, the board's instrument for isolating the model variable.
- The fixed scaffold is recent but already the most-published name. The frozen lane opened on 2025-07-20 and carries 47 of the 180 entries through 2026-02-26 — the most-published name in the extract, reached inside roughly seven months of listing activity.
- Scaffolds are vendor-crossing in practice. Ten scaffold names have published entries wearing more than one model vendor, and the fixed scaffold alone has carried ten vendors as published (case-spelling variants of one vendor merged): Anthropic, DeepSeek, Google DeepMind, Meta, Minimax, Mistral, Moonshot AI, OpenAI, Qwen, and Z.ai. For a leader this is the seam to watch: the public record shows the loop around the model is not welded to the model's maker — the same property the product layer sells as bring-your-own-model.
- Reporting discipline is arriving unevenly. No 2024 entry publishes a cost figure; the earliest cost figure in the stream is listed 2025-07-20; 32 of 112 entries in 2025 carry one; all 13 entries in 2026 do. The source's checked marker rides with 10 of 51 entries from 2024, 46 of 112 from 2025, and none of the 13 from 2026 yet. In the newest stream, every entry discloses a cost figure and none carries the checked marker — a reason to read frontier numbers as disclosure, not as audited results.
The product layer at this evidence date
Above the scaffold layer, the harness category is now visibly multi-product: the comparison matrix tracks seven shipping harnesses at its 2026-08-28 evidence date — Cursor, Kiro, OpenCode, Claude Code, Codex, Pi, and YYLO — spanning IDE-first products, terminal-first CLIs, and a control plane with no front end of its own, whose shell commands drive installed agent CLIs. Two observations about the layer's shape, each checkable against that matrix's rows. First, where a developer meets the agent still differs product to product: an IDE, a terminal, or no dedicated front end at all. Second, beneath that surface difference sits a convergence asymmetry: instructions that travel with the repository — AGENTS.md-style files, rules, steering, skills — are now common across the seven, while the records that outlive a session still differ product to product, from in-product history to exportable session files to entries bound into Git history. Instruction portability is converging on markdown; durable-record portability has not converged at all. Which harness handles which job better is the matrix's verdict to give, not this report's.
Changelog
- Edition 1 — 2026-08-29. First publication. Established the two-instrument method and admissibility rule, the five scaffold-layer findings, the two product-layer observations, and the update contract that follows. No prior edition exists; nothing is superseded.
- Erratum — 2026-08-30. Two scope sentences drafted before same-day sibling publications went stale by sequence: the versioned support and interoperability matrix and the site-wide recurring-research methodology guidance both published on 2026-08-29, this report's own evidence date, after these sections were written. Both sentences were corrected in place and now link those pages. No finding moved and no instrument changed; the evidence date stands.
Update ownership and citation
This report is a recurring asset of this site's research program — the same program that versions the leaderboard extract serving as its scaffold-layer instrument. Updates land as numbered editions on this URL. The review cycle is one quarter: a review that finds no material change records exactly that in the changelog and the evidence date stands; a review that finds an instrument has moved produces a new edition, which must re-derive every finding from the then-current instrument versions, re-date the header, and name in the changelog which findings changed and why. Adding or retiring an instrument is itself a new edition, never a silent edit.
Corrections follow an erratum rule: a factual error that does not change a finding's direction is fixed in place and logged in the changelog with the correction date; a finding whose direction changes is withdrawn by changelog entry and replaced in the next edition. The evidence date and the changelog together are this report's audit trail — there is no other.
Cite this report by its stable URL together with the edition number and the evidence date; a finding quoted without its date is being quoted as history. The leaderboard entries underneath belong to the teams that submitted them and to the project that runs the board; this report contributes the assembly, the cross-layer reading, and the maintenance contract, and invites reuse with attribution. The site-wide citation and maintenance guidance for this site's recurring research assets now lives at the recurring-research methodology reference.