RS9

Everyday agent benchmark, 30 September 2026

Results for one dated run.

These are tests based on RS9's everyday work: fixing a slow report, checking a tender, reconciling payments, picking up someone else's half-finished job. We gave 8 coding agents the same 20 tasks, one attempt each, with the same brief, files and a time limit of up to 30 minutes per task.

This is not a research study. One attempt per task means small gaps between agents mean little, and these are our tasks, not a general ranking. Read the methodology and caveats before drawing conclusions.

Latest run: . Results are a snapshot of the versions tested that day.

Overall

How each agent scored

Score out of 100

The share of tasks each agent passed outright on its single attempt. Partial credit, underneath, gives each task the share of its individual checks that passed, so close results can be told apart.

  1. 95

    Opus 5.5 with Claude Code

    19 of 20 tasks · partial credit 99 · 2 h 17 min total

  2. 85

    Sonnet 5.5 with Claude Code

    17 of 20 tasks · partial credit 99 · 1 h 35 min total

  3. 85

    GPT-6 Astra with Codex

    17 of 20 tasks · partial credit 99 · 2 h 24 min total

  4. 75

    Sol 6.1 with Codex

    15 of 20 tasks · partial credit 99 · 2 timed out · 4 h 52 min total

  5. 70

    DeepSeek V4.1 Flash with OpenCode (OpenRouter)

    14 of 20 tasks · partial credit 95 · 2 timed out · 4 h 7 min total

  6. 65

    Fable 5.1 with Claude Code

    13 of 20 tasks · partial credit 98 · 2 timed out · 3 h 17 min total

  7. 55

    Muse Spark 1.3 with Muse Code

    11 of 20 tasks · partial credit 91 · 2 timed out · 4 h 28 min total

  8. 45

    Grok 4.7 with Grok Build

    9 of 20 tasks · partial credit 76 · 10 timed out · 7 h 53 min total

4 frontier tasks were added after the top model passed all 16 of the first set; they're designed to stretch the strongest agents.

A timed-out attempt counts as not passed. Time limits run up to 30 minutes a task, and timeouts are common for some agents.

View this chart as a table
Score, tasks passed and partial credit for each agent, highest score first
ModelAgent toolScoreTasks passedPartial creditTotal time
Opus 5.5Claude Code9519 of 20992 h 17 min
Sonnet 5.5Claude Code8517 of 20991 h 35 min
GPT-6 AstraCodex8517 of 20992 h 24 min
Sol 6.1Codex7515 of 20994 h 52 min
DeepSeek V4.1 FlashOpenCode (OpenRouter)7014 of 20954 h 7 min
Fable 5.1Claude Code6513 of 20983 h 17 min
Muse Spark 1.3Muse Code5511 of 20914 h 28 min
Grok 4.7Grok Build459 of 20767 h 53 min

Over time

Is each agent holding steady?

Score, run by run

The benchmark is rerun on a regular schedule. Each line follows one agent across those dated runs, so a real change stands out from week-to-week wobble. With one attempt per task, a few points either way is normal.

Trend appears after the second run

So far there is one run, from 30 Sept. Once the benchmark has run twice, each agent gets a line here.

View this chart as a table
Score and cost per pass for each agent at each run
Model and agent tool30 Sept
Fable 5.1 with Claude Codescore 65, $5.78 per pass
Opus 5.5 with Claude Codescore 95, $1.62 per pass
Sonnet 5.5 with Claude Codescore 85, $1.00 per pass
GPT-6 Astra with Codexscore 85, $1.75 per pass
Sol 6.1 with Codexscore 75, $0.35 per pass
Grok 4.7 with Grok Buildscore 45, $3.82 per pass
Muse Spark 1.3 with Muse Codescore 55, cost n/a
DeepSeek V4.1 Flash with OpenCode (OpenRouter)score 70, $0.069 per pass

Task by task

Every result

  • Pass
  • Fail
  • Timed out
  • Error
  • Skipped
Result of each agent on each task. Time is minutes and seconds for that attempt.
TaskFable 5.1with Claude CodeOpus 5.5with Claude CodeSonnet 5.5with Claude CodeGPT-6 Astrawith CodexSol 6.1with CodexGrok 4.7with Grok BuildMuse Spark 1.3with Muse CodeDeepSeek V4.1 Flashwith OpenCode (OpenRouter)
Software
Move a billing codebase from floats to integer minor unitsFrontier
PassTime 16:37
PassTime 12:10
PassTime 7:46
PassTime 9:38
PassTime 27:35
Timed outTime 30:00
FailTime 14:38
FailTime 19:13
Fix wrong totals in a checkout service
PassTime 10:18
PassTime 6:42
PassTime 3:31
PassTime 6:27
PassTime 16:48
Timed outTime 30:00
PassTime 9:07
PassTime 12:41
Add multi-location stock with a safe data migration
PassTime 15:32
PassTime 13:46
PassTime 6:07
PassTime 12:13
PassTime 21:27
Timed outTime 30:00
FailTime 15:06
PassTime 20:29
Research and decisions
Audit a vendor's marketing claims against a pack of specs, tests and standards
PassTime 1:48
PassTime 1:09
PassTime 0:44
PassTime 1:50
PassTime 3:05
PassTime 10:34
PassTime 2:31
PassTime 2:04
Six clients, one grant: eligibility and maximum funding across amendments
PassTime 3:37
PassTime 2:28
PassTime 1:18
PassTime 2:41
PassTime 7:43
PassTime 14:28
PassTime 4:12
PassTime 3:46
Documents
Apply 14 agreed changes across a contract set
FailTime 3:60
PassTime 3:01
PassTime 2:20
PassTime 6:05
PassTime 11:04
PassTime 17:33
PassTime 7:13
PassTime 7:42
Answer a council tender from the supplier's own documents
FailTime 8:58
PassTime 3:29
FailTime 2:42
PassTime 4:33
FailTime 9:09
FailTime 30:00
FailTime 5:17
PassTime 9:17
Price a fabrication and installation tender from a messy packFrontier
FailTime 9:00
PassTime 4:39
PassTime 3:21
FailTime 4:45
FailTime 9:22
Timed outTime 30:00
FailTime 13:29
FailTime 11:47
Data
Analyse a checkout A/B test from raw event logs
PassTime 2:47
PassTime 1:45
PassTime 0:59
PassTime 4:23
PassTime 11:32
PassTime 12:37
PassTime 7:30
PassTime 2:37
October month-end: bank, ledger and payment processor
PassTime 4:41
PassTime 3:04
PassTime 2:05
FailTime 5:53
FailTime 7:49
PassTime 20:06
PassTime 29:24
PassTime 15:16
Interfaces and technical work
Build a three-page hardware shop from mockups and a behaviour specFrontier
FailTime 9:44
FailTime 8:02
PassTime 14:36
PassTime 14:35
PassTime 27:54
Timed outTime 30:00
FailTime 14:11
Timed outTime 30:00
Bill of materials and optimal cut list from a revised drawing set
PassTime 1:48
PassTime 1:09
PassTime 0:36
PassTime 1:54
PassTime 4:12
PassTime 9:31
PassTime 5:17
PassTime 2:16
Build a sheet-metal quote wizard to a 30-rule spec
FailTime 20:36
PassTime 15:27
FailTime 7:56
PassTime 14:59
Timed outTime 30:00
Timed outTime 30:00
FailTime 15:42
PassTime 21:18
Picking up where someone left off
Finish a feature after a five-day chat thread of changing requirements
PassTime 4:42
PassTime 3:37
FailTime 2:37
PassTime 10:02
PassTime 12:52
PassTime 20:49
PassTime 8:30
PassTime 5:15
Diagnose a two-cause checkout incident across three services
PassTime 4:54
PassTime 3:01
PassTime 2:38
PassTime 3:32
PassTime 7:32
PassTime 15:14
PassTime 5:41
PassTime 6:01
CAD and fabrication
Parametric sensor-board enclosure in OpenSCAD
PassTime 5:39
PassTime 10:49
PassTime 3:10
PassTime 7:14
PassTime 11:12
Timed outTime 30:00
PassTime 27:38
PassTime 11:16
Laser-cut flat pattern for a formed bracket
PassTime 1:44
PassTime 1:21
PassTime 1:05
PassTime 3:31
PassTime 6:06
PassTime 22:35
FailTime 10:53
PassTime 11:55
Hinged sensor clamp and wall bracket, measured from a vendor STLFrontier
Timed outTime 30:00
PassTime 13:55
PassTime 10:23
PassTime 10:22
PassTime 16:53
Timed outTime 30:00
Timed outTime 30:00
Timed outTime 30:00
Front-end design
Design a landing page that borrows an inspiration image's visual language
PassTime 10:32
PassTime 6:59
PassTime 3:05
PassTime 6:16
PassTime 19:38
Timed outTime 30:00
PassTime 11:45
FailTime 8:53
Rebuild a landing page from desktop, tablet and mobile mockups
Timed outTime 30:00
PassTime 20:26
PassTime 17:51
FailTime 13:36
Timed outTime 30:00
Timed outTime 30:00
Timed outTime 30:00
FailTime 15:01

Speed

Passing versus time taken

Score against time taken

Each dot is one model with its agent tool. Higher means a better score; further left means less total time spent across its 20 attempts. The best spot is the top left. Hover or tab to a dot to see which one it is.

Score against total timeScatter chart. Each dot is one model with its agent tool. Higher is a better score; the score axis starts at 40. Fable 5.1 with Claude Code: score 65 (13 of 20 tasks), 3 h 17 m total; Opus 5.5 with Claude Code: score 95 (19 of 20 tasks), 2 h 17 m total; Sonnet 5.5 with Claude Code: score 85 (17 of 20 tasks), 1 h 35 m total; GPT-6 Astra with Codex: score 85 (17 of 20 tasks), 2 h 24 m total; Sol 6.1 with Codex: score 75 (15 of 20 tasks), 4 h 52 m total; Grok 4.7 with Grok Build: score 45 (9 of 20 tasks), 7 h 53 m total; Muse Spark 1.3 with Muse Code: score 55 (11 of 20 tasks), 4 h 28 m total; DeepSeek V4.1 Flash with OpenCode (OpenRouter): score 70 (14 of 20 tasks), 4 h 7 m total.4050607080901000 min2 h4 h6 h8 hTotal time across all tasks (shorter is faster)Score (axis starts at 40)Sonnet 5.5Opus 5.5GPT-6 AstraFable 5.1DeepSeek V4.1 FlashMuse Spark 1.3Sol 6.1Grok 4.7

Times include each provider's server speed and queueing.

View this chart as a table
Score and total time for each agent
ModelAgent toolScoreTasks passedTotal time
Opus 5.5Claude Code9519 of 202 h 17 min
Sonnet 5.5Claude Code8517 of 201 h 35 min
GPT-6 AstraCodex8517 of 202 h 24 min
Sol 6.1Codex7515 of 204 h 52 min
DeepSeek V4.1 FlashOpenCode (OpenRouter)7014 of 204 h 7 min
Fable 5.1Claude Code6513 of 203 h 17 min
Muse Spark 1.3Muse Code5511 of 204 h 28 min
Grok 4.7Grok Build459 of 207 h 53 min

Cost

What each passed task cost

Subscription agents don't bill per task. Their cost here is what the same tokens would cost at the provider's public API price: we count the tokens each attempt used and price them at list rates, so every agent can be compared on one scale.

A model run through OpenRouter shows what we were actually charged, which can be well below list price when the host discounts it. If we can't find a public price for a model, it says “price unknown” rather than $0.

Cost per task passed

What each agent's spend works out to for every task it got right: total cost divided by tasks passed. Shorter bar means cheaper. All amounts are US dollars.

  1. DeepSeek V4.1 Flash with OpenCode (OpenRouter)

    $0.069 per pass

    $0.96 total · Real charge via OpenRouter

  2. Sonnet 5.5 with Claude Code

    $1.00 per pass

    $16.93 total · Subscription, at API list price

  3. Opus 5.5 with Claude Code

    $1.62 per pass

    $30.77 total · Subscription, at API list price

  4. GPT-6 Astra with Codex

    $1.75 per pass

    $29.79 total · Subscription, at API list price

  5. Grok 4.7 with Grok Build

    $3.82 per pass

    $34.41 total · Subscription, at API list price, a slight undercount: 11 attempts were cut off before final usage was reported

  6. Fable 5.1 with Claude Code

    $5.78 per pass

    $75.18 total · Subscription, at API list price

  7. Muse Spark 1.3 with Muse Code

    price unknown

    No public API price

  8. Sol 6.1 with Codex

    price unknown

    Some attempts have no usage data, so no total

View this chart as a table
Cost per task passed, cheapest first
ModelAgent toolCost per passTotal costTasks passedBasis
DeepSeek V4.1 FlashOpenCode (OpenRouter)$0.069$0.9614 of 20What OpenRouter actually charged
Sonnet 5.5Claude Code$1.00$16.9317 of 20Subscription: what the same tokens would cost at the provider's public API price
Opus 5.5Claude Code$1.62$30.7719 of 20Subscription: what the same tokens would cost at the provider's public API price
GPT-6 AstraCodex$1.75$29.7917 of 20Subscription: what the same tokens would cost at the provider's public API price
Grok 4.7Grok Build$3.82$34.419 of 20Subscription: what the same tokens would cost at the provider's public API price
Fable 5.1Claude Code$5.78$75.1813 of 20Subscription: what the same tokens would cost at the provider's public API price
Muse Spark 1.3Muse Codeprice unknownprice unknown11 of 20No public API price, so no cost is shown
Sol 6.1Codexprice unknownprice unknown15 of 20Some attempts reported no usage, so no total is shown

Score against cost per task passed

Each dot is one model with its agent tool. Higher means a better score; further left means each passed task cost less. The best spot is the top left. Hover or tab to a dot to see which one it is.

Score against cost per task passedScatter chart. Each dot is one model with its agent tool. Higher is a better score; the score axis starts at 40. Fable 5.1 with Claude Code: score 65 (13 of 20 tasks), $5.78 per pass; Opus 5.5 with Claude Code: score 95 (19 of 20 tasks), $1.62 per pass; Sonnet 5.5 with Claude Code: score 85 (17 of 20 tasks), $1.00 per pass; GPT-6 Astra with Codex: score 85 (17 of 20 tasks), $1.75 per pass; Sol 6.1 with Codex: score 75 (15 of 20 tasks), $0.35 per pass; Grok 4.7 with Grok Build: score 45 (9 of 20 tasks), $3.82 per pass; DeepSeek V4.1 Flash with OpenCode (OpenRouter): score 70 (14 of 20 tasks), $0.07 per pass.405060708090100$0$2$4$6$8Cost per task passed, US dollars (further left is cheaper)Score (axis starts at 40)DeepSeek V4.1 FlashSol 6.1Sonnet 5.5Opus 5.5GPT-6 AstraGrok 4.7Fable 5.1

Not plotted: Muse Spark 1.3 with Muse Code (price unknown).

View this chart as a table
Score and cost per task passed for each agent
ModelAgent toolScoreTasks passedCost per pass
Opus 5.5Claude Code9519 of 20$1.62
Sonnet 5.5Claude Code8517 of 20$1.00
GPT-6 AstraCodex8517 of 20$1.75
Sol 6.1Codex7515 of 20$0.35
DeepSeek V4.1 FlashOpenCode (OpenRouter)7014 of 20$0.069
Fable 5.1Claude Code6513 of 20$5.78
Muse Spark 1.3Muse Code5511 of 20price unknown
Grok 4.7Grok Build459 of 20$3.82

The agents

What we tested

Times include each provider's server speed and queueing.

  • Fable 5.1

    with Claude Code

    Model
    Claude Fable 5.1
    Reasoning
    High
    Client version
    2.1.285 (Claude Code)
    Score
    65 · partial credit 98
    Passed
    13 of 20
    Timed out
    2 of 20
    Total time
    3 h 17 min
    Total cost
    US$75.18
    Cost per pass
    US$5.78

    Subscription, at API list price.

    Claude Code in headless mode, high reasoning effort, on a Claude subscription.

  • Opus 5.5

    with Claude Code

    Model
    Claude Opus 5.5
    Reasoning
    High
    Client version
    2.1.285 (Claude Code)
    Score
    95 · partial credit 99
    Passed
    19 of 20
    Total time
    2 h 17 min
    Total cost
    US$30.77
    Cost per pass
    US$1.62

    Subscription, at API list price.

    Claude Code in headless mode, high reasoning effort, on a Claude subscription.

  • Sonnet 5.5

    with Claude Code

    Model
    Claude Sonnet 5.5
    Reasoning
    High
    Client version
    2.1.285 (Claude Code)
    Score
    85 · partial credit 99
    Passed
    17 of 20
    Total time
    1 h 35 min
    Total cost
    US$16.93
    Cost per pass
    US$1.00

    Subscription, at API list price.

    Claude Code in headless mode, high reasoning effort, on a Claude subscription.

  • GPT-6 Astra

    with Codex

    Model
    GPT-6 Astra
    Reasoning
    High
    Client version
    codex-cli 0.159.2
    Score
    85 · partial credit 99
    Passed
    17 of 20
    Total time
    2 h 24 min
    Total cost
    US$29.79
    Cost per pass
    US$1.75

    Subscription, at API list price.

    Codex CLI (exec mode), high reasoning effort, on a ChatGPT subscription.

  • Sol 6.1

    with Codex

    Model
    ChatGPT Sol 6.1
    Reasoning
    High
    Client version
    codex-cli 0.159.2
    Score
    75 · partial credit 99
    Passed
    15 of 20
    Timed out
    2 of 20
    Total time
    4 h 52 min
    Total cost
    price unknown
    Cost per pass
    price unknown

    Some attempts have no usage data, so no total.

    Codex CLI (exec mode), high reasoning effort, on a ChatGPT subscription.

  • Grok 4.7

    with Grok Build

    Model
    Grok 4.7
    Reasoning
    High
    Client version
    grok 1.0.44 (5b807183dd79)
    Score
    45 · partial credit 76
    Passed
    9 of 20
    Timed out
    10 of 20
    Total time
    7 h 53 min
    Total cost
    US$34.41
    Cost per pass
    US$3.82

    Subscription, at API list price. Slight undercount: 11 attempts were cut off before final usage was reported.

    Grok Build CLI in headless mode, high reasoning effort, on a SuperGrok subscription.

  • Muse Spark 1.3

    with Muse Code

    Model
    Muse Spark 1.3
    Reasoning
    High
    Client version
    Muse Code 1.4.1 (1.4.1-R4503.1)
    Score
    55 · partial credit 91
    Passed
    11 of 20
    Timed out
    2 of 20
    Total time
    4 h 28 min
    Total cost
    price unknown
    Cost per pass
    price unknown

    No public API price.

    Muse Code CLI (exec mode), high reasoning effort, on a Muse Code subscription. Meta publishes no per-token price, so cost is unknown.

  • DeepSeek V4.1 Flash

    with OpenCode (OpenRouter)

    Model
    DeepSeek V4.1 Flash
    Reasoning
    High
    Client version
    1.18.33
    Score
    70 · partial credit 95
    Passed
    14 of 20
    Timed out
    2 of 20
    Total time
    4 h 7 min
    Total cost
    US$0.96
    Cost per pass
    US$0.069

    Real charge via OpenRouter. At list price it would be US$0.20 per pass.

    OpenCode via OpenRouter, served by DeepInfra (US), high reasoning effort. Cost is the real charge.

Check our numbers

Downloads

History

Runs

  • 30 September 2026 · this page

There are no earlier runs yet.