These are tests based on RS9's everyday work: fixing a slow report, checking a tender, reconciling payments, picking up someone else's half-finished job. We gave 8 coding agents the same 20 tasks, one attempt each, with the same brief, files and a time limit of up to 30 minutes per task.
This is not a research study. One attempt per task means small gaps between agents mean little, and these are our tasks, not a general ranking. Read the methodology and caveats before drawing conclusions.
Latest run: . Results are a snapshot of the versions tested that day.
Overall
How each agent scored
Score out of 100
The share of tasks each agent passed outright on its single attempt. Partial credit, underneath, gives each task the share of its individual checks that passed, so close results can be told apart.
4 frontier tasks were added after the top model passed all 16 of the first set; they're designed to stretch the strongest agents.
A timed-out attempt counts as not passed. Time limits run up to 30 minutes a task, and timeouts are common for some agents.
View this chart as a table
| Model | Agent tool | Score | Tasks passed | Partial credit | Total time |
|---|---|---|---|---|---|
| Opus 5.5 | Claude Code | 95 | 19 of 20 | 99 | 2 h 17 min |
| Sonnet 5.5 | Claude Code | 85 | 17 of 20 | 99 | 1 h 35 min |
| GPT-6 Astra | Codex | 85 | 17 of 20 | 99 | 2 h 24 min |
| Sol 6.1 | Codex | 75 | 15 of 20 | 99 | 4 h 52 min |
| DeepSeek V4.1 Flash | OpenCode (OpenRouter) | 70 | 14 of 20 | 95 | 4 h 7 min |
| Fable 5.1 | Claude Code | 65 | 13 of 20 | 98 | 3 h 17 min |
| Muse Spark 1.3 | Muse Code | 55 | 11 of 20 | 91 | 4 h 28 min |
| Grok 4.7 | Grok Build | 45 | 9 of 20 | 76 | 7 h 53 min |
Over time
Is each agent holding steady?
Score, run by run
The benchmark is rerun on a regular schedule. Each line follows one agent across those dated runs, so a real change stands out from week-to-week wobble. With one attempt per task, a few points either way is normal.
Trend appears after the second run
So far there is one run, from 30 Sept. Once the benchmark has run twice, each agent gets a line here.
View this chart as a table
| Model and agent tool | 30 Sept |
|---|---|
| Fable 5.1 with Claude Code | score 65, $5.78 per pass |
| Opus 5.5 with Claude Code | score 95, $1.62 per pass |
| Sonnet 5.5 with Claude Code | score 85, $1.00 per pass |
| GPT-6 Astra with Codex | score 85, $1.75 per pass |
| Sol 6.1 with Codex | score 75, $0.35 per pass |
| Grok 4.7 with Grok Build | score 45, $3.82 per pass |
| Muse Spark 1.3 with Muse Code | score 55, cost n/a |
| DeepSeek V4.1 Flash with OpenCode (OpenRouter) | score 70, $0.069 per pass |
Task by task
Every result
- Pass
- Fail
- Timed out
- Error
- Skipped
| Task | Fable 5.1with Claude Code | Opus 5.5with Claude Code | Sonnet 5.5with Claude Code | GPT-6 Astrawith Codex | Sol 6.1with Codex | Grok 4.7with Grok Build | Muse Spark 1.3with Muse Code | DeepSeek V4.1 Flashwith OpenCode (OpenRouter) |
|---|---|---|---|---|---|---|---|---|
| Software | ||||||||
| Move a billing codebase from floats to integer minor unitsFrontier | PassTime 16:37 | PassTime 12:10 | PassTime 7:46 | PassTime 9:38 | PassTime 27:35 | Timed outTime 30:00 | FailTime 14:38 | FailTime 19:13 |
| Fix wrong totals in a checkout service | PassTime 10:18 | PassTime 6:42 | PassTime 3:31 | PassTime 6:27 | PassTime 16:48 | Timed outTime 30:00 | PassTime 9:07 | PassTime 12:41 |
| Add multi-location stock with a safe data migration | PassTime 15:32 | PassTime 13:46 | PassTime 6:07 | PassTime 12:13 | PassTime 21:27 | Timed outTime 30:00 | FailTime 15:06 | PassTime 20:29 |
| Research and decisions | ||||||||
| Audit a vendor's marketing claims against a pack of specs, tests and standards | PassTime 1:48 | PassTime 1:09 | PassTime 0:44 | PassTime 1:50 | PassTime 3:05 | PassTime 10:34 | PassTime 2:31 | PassTime 2:04 |
| Six clients, one grant: eligibility and maximum funding across amendments | PassTime 3:37 | PassTime 2:28 | PassTime 1:18 | PassTime 2:41 | PassTime 7:43 | PassTime 14:28 | PassTime 4:12 | PassTime 3:46 |
| Documents | ||||||||
| Apply 14 agreed changes across a contract set | FailTime 3:60 | PassTime 3:01 | PassTime 2:20 | PassTime 6:05 | PassTime 11:04 | PassTime 17:33 | PassTime 7:13 | PassTime 7:42 |
| Answer a council tender from the supplier's own documents | FailTime 8:58 | PassTime 3:29 | FailTime 2:42 | PassTime 4:33 | FailTime 9:09 | FailTime 30:00 | FailTime 5:17 | PassTime 9:17 |
| Price a fabrication and installation tender from a messy packFrontier | FailTime 9:00 | PassTime 4:39 | PassTime 3:21 | FailTime 4:45 | FailTime 9:22 | Timed outTime 30:00 | FailTime 13:29 | FailTime 11:47 |
| Data | ||||||||
| Analyse a checkout A/B test from raw event logs | PassTime 2:47 | PassTime 1:45 | PassTime 0:59 | PassTime 4:23 | PassTime 11:32 | PassTime 12:37 | PassTime 7:30 | PassTime 2:37 |
| October month-end: bank, ledger and payment processor | PassTime 4:41 | PassTime 3:04 | PassTime 2:05 | FailTime 5:53 | FailTime 7:49 | PassTime 20:06 | PassTime 29:24 | PassTime 15:16 |
| Interfaces and technical work | ||||||||
| Build a three-page hardware shop from mockups and a behaviour specFrontier | FailTime 9:44 | FailTime 8:02 | PassTime 14:36 | PassTime 14:35 | PassTime 27:54 | Timed outTime 30:00 | FailTime 14:11 | Timed outTime 30:00 |
| Bill of materials and optimal cut list from a revised drawing set | PassTime 1:48 | PassTime 1:09 | PassTime 0:36 | PassTime 1:54 | PassTime 4:12 | PassTime 9:31 | PassTime 5:17 | PassTime 2:16 |
| Build a sheet-metal quote wizard to a 30-rule spec | FailTime 20:36 | PassTime 15:27 | FailTime 7:56 | PassTime 14:59 | Timed outTime 30:00 | Timed outTime 30:00 | FailTime 15:42 | PassTime 21:18 |
| Picking up where someone left off | ||||||||
| Finish a feature after a five-day chat thread of changing requirements | PassTime 4:42 | PassTime 3:37 | FailTime 2:37 | PassTime 10:02 | PassTime 12:52 | PassTime 20:49 | PassTime 8:30 | PassTime 5:15 |
| Diagnose a two-cause checkout incident across three services | PassTime 4:54 | PassTime 3:01 | PassTime 2:38 | PassTime 3:32 | PassTime 7:32 | PassTime 15:14 | PassTime 5:41 | PassTime 6:01 |
| CAD and fabrication | ||||||||
| Parametric sensor-board enclosure in OpenSCAD | PassTime 5:39 | PassTime 10:49 | PassTime 3:10 | PassTime 7:14 | PassTime 11:12 | Timed outTime 30:00 | PassTime 27:38 | PassTime 11:16 |
| Laser-cut flat pattern for a formed bracket | PassTime 1:44 | PassTime 1:21 | PassTime 1:05 | PassTime 3:31 | PassTime 6:06 | PassTime 22:35 | FailTime 10:53 | PassTime 11:55 |
| Hinged sensor clamp and wall bracket, measured from a vendor STLFrontier | Timed outTime 30:00 | PassTime 13:55 | PassTime 10:23 | PassTime 10:22 | PassTime 16:53 | Timed outTime 30:00 | Timed outTime 30:00 | Timed outTime 30:00 |
| Front-end design | ||||||||
| Design a landing page that borrows an inspiration image's visual language | PassTime 10:32 | PassTime 6:59 | PassTime 3:05 | PassTime 6:16 | PassTime 19:38 | Timed outTime 30:00 | PassTime 11:45 | FailTime 8:53 |
| Rebuild a landing page from desktop, tablet and mobile mockups | Timed outTime 30:00 | PassTime 20:26 | PassTime 17:51 | FailTime 13:36 | Timed outTime 30:00 | Timed outTime 30:00 | Timed outTime 30:00 | FailTime 15:01 |
Speed
Passing versus time taken
Score against time taken
Each dot is one model with its agent tool. Higher means a better score; further left means less total time spent across its 20 attempts. The best spot is the top left. Hover or tab to a dot to see which one it is.
Times include each provider's server speed and queueing.
View this chart as a table
| Model | Agent tool | Score | Tasks passed | Total time |
|---|---|---|---|---|
| Opus 5.5 | Claude Code | 95 | 19 of 20 | 2 h 17 min |
| Sonnet 5.5 | Claude Code | 85 | 17 of 20 | 1 h 35 min |
| GPT-6 Astra | Codex | 85 | 17 of 20 | 2 h 24 min |
| Sol 6.1 | Codex | 75 | 15 of 20 | 4 h 52 min |
| DeepSeek V4.1 Flash | OpenCode (OpenRouter) | 70 | 14 of 20 | 4 h 7 min |
| Fable 5.1 | Claude Code | 65 | 13 of 20 | 3 h 17 min |
| Muse Spark 1.3 | Muse Code | 55 | 11 of 20 | 4 h 28 min |
| Grok 4.7 | Grok Build | 45 | 9 of 20 | 7 h 53 min |
Cost
What each passed task cost
Subscription agents don't bill per task. Their cost here is what the same tokens would cost at the provider's public API price: we count the tokens each attempt used and price them at list rates, so every agent can be compared on one scale.
A model run through OpenRouter shows what we were actually charged, which can be well below list price when the host discounts it. If we can't find a public price for a model, it says “price unknown” rather than $0.
Cost per task passed
What each agent's spend works out to for every task it got right: total cost divided by tasks passed. Shorter bar means cheaper. All amounts are US dollars.
View this chart as a table
| Model | Agent tool | Cost per pass | Total cost | Tasks passed | Basis |
|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | OpenCode (OpenRouter) | $0.069 | $0.96 | 14 of 20 | What OpenRouter actually charged |
| Sonnet 5.5 | Claude Code | $1.00 | $16.93 | 17 of 20 | Subscription: what the same tokens would cost at the provider's public API price |
| Opus 5.5 | Claude Code | $1.62 | $30.77 | 19 of 20 | Subscription: what the same tokens would cost at the provider's public API price |
| GPT-6 Astra | Codex | $1.75 | $29.79 | 17 of 20 | Subscription: what the same tokens would cost at the provider's public API price |
| Grok 4.7 | Grok Build | $3.82 | $34.41 | 9 of 20 | Subscription: what the same tokens would cost at the provider's public API price |
| Fable 5.1 | Claude Code | $5.78 | $75.18 | 13 of 20 | Subscription: what the same tokens would cost at the provider's public API price |
| Muse Spark 1.3 | Muse Code | price unknown | price unknown | 11 of 20 | No public API price, so no cost is shown |
| Sol 6.1 | Codex | price unknown | price unknown | 15 of 20 | Some attempts reported no usage, so no total is shown |
Score against cost per task passed
Each dot is one model with its agent tool. Higher means a better score; further left means each passed task cost less. The best spot is the top left. Hover or tab to a dot to see which one it is.
Not plotted: Muse Spark 1.3 with Muse Code (price unknown).
View this chart as a table
| Model | Agent tool | Score | Tasks passed | Cost per pass |
|---|---|---|---|---|
| Opus 5.5 | Claude Code | 95 | 19 of 20 | $1.62 |
| Sonnet 5.5 | Claude Code | 85 | 17 of 20 | $1.00 |
| GPT-6 Astra | Codex | 85 | 17 of 20 | $1.75 |
| Sol 6.1 | Codex | 75 | 15 of 20 | $0.35 |
| DeepSeek V4.1 Flash | OpenCode (OpenRouter) | 70 | 14 of 20 | $0.069 |
| Fable 5.1 | Claude Code | 65 | 13 of 20 | $5.78 |
| Muse Spark 1.3 | Muse Code | 55 | 11 of 20 | price unknown |
| Grok 4.7 | Grok Build | 45 | 9 of 20 | $3.82 |
The agents
What we tested
Times include each provider's server speed and queueing.
Fable 5.1
with Claude Code
- Model
- Claude Fable 5.1
- Reasoning
- High
- Client version
- 2.1.285 (Claude Code)
- Score
- 65 · partial credit 98
- Passed
- 13 of 20
- Timed out
- 2 of 20
- Total time
- 3 h 17 min
- Total cost
- US$75.18
- Cost per pass
- US$5.78
Subscription, at API list price.
Claude Code in headless mode, high reasoning effort, on a Claude subscription.
Opus 5.5
with Claude Code
- Model
- Claude Opus 5.5
- Reasoning
- High
- Client version
- 2.1.285 (Claude Code)
- Score
- 95 · partial credit 99
- Passed
- 19 of 20
- Total time
- 2 h 17 min
- Total cost
- US$30.77
- Cost per pass
- US$1.62
Subscription, at API list price.
Claude Code in headless mode, high reasoning effort, on a Claude subscription.
Sonnet 5.5
with Claude Code
- Model
- Claude Sonnet 5.5
- Reasoning
- High
- Client version
- 2.1.285 (Claude Code)
- Score
- 85 · partial credit 99
- Passed
- 17 of 20
- Total time
- 1 h 35 min
- Total cost
- US$16.93
- Cost per pass
- US$1.00
Subscription, at API list price.
Claude Code in headless mode, high reasoning effort, on a Claude subscription.
GPT-6 Astra
with Codex
- Model
- GPT-6 Astra
- Reasoning
- High
- Client version
- codex-cli 0.159.2
- Score
- 85 · partial credit 99
- Passed
- 17 of 20
- Total time
- 2 h 24 min
- Total cost
- US$29.79
- Cost per pass
- US$1.75
Subscription, at API list price.
Codex CLI (exec mode), high reasoning effort, on a ChatGPT subscription.
Sol 6.1
with Codex
- Model
- ChatGPT Sol 6.1
- Reasoning
- High
- Client version
- codex-cli 0.159.2
- Score
- 75 · partial credit 99
- Passed
- 15 of 20
- Timed out
- 2 of 20
- Total time
- 4 h 52 min
- Total cost
- price unknown
- Cost per pass
- price unknown
Some attempts have no usage data, so no total.
Codex CLI (exec mode), high reasoning effort, on a ChatGPT subscription.
Grok 4.7
with Grok Build
- Model
- Grok 4.7
- Reasoning
- High
- Client version
- grok 1.0.44 (5b807183dd79)
- Score
- 45 · partial credit 76
- Passed
- 9 of 20
- Timed out
- 10 of 20
- Total time
- 7 h 53 min
- Total cost
- US$34.41
- Cost per pass
- US$3.82
Subscription, at API list price. Slight undercount: 11 attempts were cut off before final usage was reported.
Grok Build CLI in headless mode, high reasoning effort, on a SuperGrok subscription.
Muse Spark 1.3
with Muse Code
- Model
- Muse Spark 1.3
- Reasoning
- High
- Client version
- Muse Code 1.4.1 (1.4.1-R4503.1)
- Score
- 55 · partial credit 91
- Passed
- 11 of 20
- Timed out
- 2 of 20
- Total time
- 4 h 28 min
- Total cost
- price unknown
- Cost per pass
- price unknown
No public API price.
Muse Code CLI (exec mode), high reasoning effort, on a Muse Code subscription. Meta publishes no per-token price, so cost is unknown.
DeepSeek V4.1 Flash
with OpenCode (OpenRouter)
- Model
- DeepSeek V4.1 Flash
- Reasoning
- High
- Client version
- 1.18.33
- Score
- 70 · partial credit 95
- Passed
- 14 of 20
- Timed out
- 2 of 20
- Total time
- 4 h 7 min
- Total cost
- US$0.96
- Cost per pass
- US$0.069
Real charge via OpenRouter. At list price it would be US$0.20 per pass.
OpenCode via OpenRouter, served by DeepInfra (US), high reasoning effort. Cost is the real charge.
Check our numbers
Downloads
History
Runs
- 30 September 2026 · this page
There are no earlier runs yet.