RS9

RS9 Everyday Agent Bench, 2 October 2026

One dated run. These are the results as they stood that day.

RS9 Everyday Agent Bench

RS9

Best results

Score out of 100, highest first

  1. Opus 5.5 with Claude Code

    95

  2. Sonnet 5.5 with Claude Code

    75

  3. Fable 5.1 with Claude Code

    70

Best value

Cost per task passed, cheapest first

  1. Sonnet 5.5 with Claude Code

    $1.11per task passed

  2. Opus 5.5 with Claude Code

    $1.28per task passed

  3. Fable 5.1 with Claude Code

    $5.12per task passed

Subscription agents: what the same work would cost at public API prices. OpenRouter: the real charge.

20 tasks · one attempt each

rs9.com.au/benchmarks

Score is the percentage of tasks passed, one attempt each. Partial credit (in the table below) gives each task the share of its checks that passed, so it can split agents on the same score.

Value is total cost divided by tasks passed, so an agent that's cheap but fails a lot doesn't look good. It sits beside the score and never changes it.

A few of the 20 tasks are “frontier” tasks, added to stretch the top models.

View the results as a table
Score, tasks passed and partial credit for each agent, highest score first
ModelAgent toolScoreTasks passedPartial creditTotal time
Opus 5.5Claude Code9519 of 20991 h 48 min
Sonnet 5.5Claude Code7515 of 20991 h 36 min
Fable 5.1Claude Code7014 of 20992 h 42 min
View the value board as a table
Cost per task passed for each agent, cheapest first. Subscription agents show the public API price equivalent. OpenRouter shows the real charge.
ModelAgent toolCost per task passedTasks passed
Sonnet 5.5Claude Code$1.11 per task passed15 of 20
Opus 5.5Claude Code$1.28 per task passed19 of 20
Fable 5.1Claude Code$5.12 per task passed14 of 20

We gave 3 coding agents the same 20 tasks from the kind of work RS9 does, then checked whether each one got the job done. One attempt each, same brief and files, 60 minutes a task.

It's not a research study. Each agent gets one go at each task, so a gap of a task or two means very little. And these are our tasks, so how an agent does here says little about how it'll do on yours. How it's scored, and where it falls short, is on the methodology page.

Latest run: . These are the versions we tested that day. Update an agent and its numbers can change.

Over time

Is each agent holding steady?

Score, run by run

We rerun the benchmark regularly. Each line follows one agent across the dated runs. With one attempt per task, a few points either way is normal wobble. A steady drift is what to look for.

Select an agent in the key to pick out its line.

Score over timeLine chart with one line per agent across 2 runs from 30 September 2026 to 2 October 2026. The same numbers are in the table below the chart.025507510030 Sept2 OctScore (out of 100)Run date

A diamond means the agent's model or client version changed since its last run. Select an agent to step through its points with the keyboard.

View this chart as a table
Score and cost per pass for each agent at each run
Model and agent tool30 Sept2 Oct
Fable 5.1 with Claude Codescore 70, $5.91 per passscore 70, $5.12 per pass
Opus 5.5 with Claude Codescore 95, $1.62 per passscore 95, $1.28 per pass
Sonnet 5.5 with Claude Codescore 85, $1.00 per passscore 75, $1.11 per pass
GPT-6 Astra with Codexscore 85, $1.75 per passnot run
Sol 6.1 with Codexscore 80, $0.45 per passnot run
Grok 4.7 with Grok Buildscore 65, $4.56 per passnot run
Muse Spark 1.3 with Muse Codescore 55, cost n/anot run
DeepSeek V4.1 Flash with OpenCode (OpenRouter)score 70, $0.085 per passnot run

Task by task

Every result

  • Pass
  • Fail
  • Timed out
  • Error
  • Skipped
Result of each agent on each task. Time is minutes and seconds for that attempt.
TaskFable 5.1with Claude CodeOpus 5.5with Claude CodeSonnet 5.5with Claude Code
Software
Codebase-wide changeFrontier
FailTime 10:43
PassTime 9:09
PassTime 6:33
Checkout bug fixing
PassTime 8:58
PassTime 4:45
PassTime 4:15
Database migration
PassTime 7:39
PassTime 6:09
PassTime 4:40
Research and decisions
Fact-check product claims
PassTime 2:51
PassTime 1:10
PassTime 0:39
Grant eligibility research
PassTime 5:28
PassTime 2:25
FailTime 1:10
Documents
Contract revisions
PassTime 4:20
PassTime 2:32
PassTime 2:08
Tender response
PassTime 6:03
PassTime 3:44
PassTime 2:06
Tender pricingFrontier
FailTime 7:49
PassTime 4:28
PassTime 4:06
Data
A/B test analysis
PassTime 2:32
PassTime 1:39
PassTime 0:59
Month-end reconciliation
PassTime 3:17
PassTime 2:45
PassTime 2:08
Interfaces and technical work
Small web shopFrontier
FailTime 13:57
PassTime 10:32
FailTime 12:35
Fabrication cut list
PassTime 1:37
PassTime 0:58
PassTime 5:21
Quote form with pricing rules
FailTime 10:15
PassTime 15:29
FailTime 5:08
Picking up where someone left off
Pick up someone else's half-finished feature
PassTime 5:21
PassTime 3:50
PassTime 2:13
Incident investigation
PassTime 6:06
PassTime 3:02
PassTime 2:28
CAD and fabrication
Electronics enclosure (3D CAD)
PassTime 5:34
PassTime 4:12
PassTime 1:30
Sheet-metal part (2D CAD)
FailTime 1:59
PassTime 1:14
FailTime 0:42
Multi-part CAD assemblyFrontier
PassTime 19:47
PassTime 12:24
PassTime 14:16
Front-end design
Page design from a reference
PassTime 8:11
PassTime 4:46
FailTime 2:24
Build a page from a mockup
FailTime 29:42
FailTime 12:54
PassTime 21:02

Speed

Passing versus time taken

Score against time taken

Each dot is one model with its agent tool. Higher is a better score. Further left means less total time across its 20 attempts. Top left is the best spot. Hover or tab to a dot to see which is which.

Score against total timeScatter chart. Each dot is one model with its agent tool. Higher is a better score; the score axis starts at 60. Fable 5.1 with Claude Code: score 70 (14 of 20 tasks), 2 h 42 m total; Opus 5.5 with Claude Code: score 95 (19 of 20 tasks), 1 h 48 m total; Sonnet 5.5 with Claude Code: score 75 (15 of 20 tasks), 1 h 36 m total.607080901000 min45 min1 h 30 m2 h 15 m3 hTotal time across all tasks (shorter is faster)Score (axis starts at 60)Sonnet 5.5Opus 5.5Fable 5.1

Times include the provider's server speed and queueing, so a slow time can mean a busy server.

View this chart as a table
Score and total time for each agent
ModelAgent toolScoreTasks passedTotal time
Opus 5.5Claude Code9519 of 201 h 48 min
Sonnet 5.5Claude Code7515 of 201 h 36 min
Fable 5.1Claude Code7014 of 202 h 42 min

Cost

What each passed task cost

Subscription agents don't bill per task. So we count the tokens each attempt used and price them at the provider's public API rates. It's not what anyone paid, but it puts every agent on one scale.

The OpenRouter model shows the real charge, which can sit well below list price if the host discounts it. If there's no public price for a model, it says “price unknown”. Never $0.

Cost per task passed

Total cost divided by tasks passed, for each agent. A shorter bar is cheaper. All amounts are US dollars.

  1. Sonnet 5.5 with Claude Code

    $1.11 per pass

    $16.63 total · Subscription, at API list price

  2. Opus 5.5 with Claude Code

    $1.28 per pass

    $24.25 total · Subscription, at API list price

  3. Fable 5.1 with Claude Code

    $5.12 per pass

    $71.66 total · Subscription, at API list price

View this chart as a table
Cost per task passed, cheapest first
ModelAgent toolCost per passTotal costTasks passedBasis
Sonnet 5.5Claude Code$1.11$16.6315 of 20Subscription: what the same tokens would cost at the provider's public API price
Opus 5.5Claude Code$1.28$24.2519 of 20Subscription: what the same tokens would cost at the provider's public API price
Fable 5.1Claude Code$5.12$71.6614 of 20Subscription: what the same tokens would cost at the provider's public API price

Score against cost per task passed

Each dot is one model with its agent tool. Higher is a better score. Further left means each passed task cost less. Top left is the best spot. Hover or tab to a dot to see which is which.

Score against cost per task passedScatter chart. Each dot is one model with its agent tool. Higher is a better score; the score axis starts at 60. Fable 5.1 with Claude Code: score 70 (14 of 20 tasks), $5.12 per pass; Opus 5.5 with Claude Code: score 95 (19 of 20 tasks), $1.28 per pass; Sonnet 5.5 with Claude Code: score 75 (15 of 20 tasks), $1.11 per pass.60708090100$0$2$4$6$8Cost per task passed, US dollars (further left is cheaper)Score (axis starts at 60)Sonnet 5.5Opus 5.5Fable 5.1
View this chart as a table
Score and cost per task passed for each agent
ModelAgent toolScoreTasks passedCost per pass
Opus 5.5Claude Code9519 of 20$1.28
Sonnet 5.5Claude Code7515 of 20$1.11
Fable 5.1Claude Code7014 of 20$5.12

The agents

What we tested

Times include the provider's server speed and queueing, so a slow time can mean a busy server.

  • Fable 5.1

    with Claude Code

    Model
    Claude Fable 5.1
    Reasoning
    High
    Client version
    2.1.285 (Claude Code)
    Score
    70 · partial credit 99
    Passed
    14 of 20
    Total time
    2 h 42 min
    Total cost
    US$71.66
    Cost per pass
    US$5.12

    Subscription, at API list price.

    Claude Code in headless mode, high reasoning effort, on a Claude subscription.

  • Opus 5.5

    with Claude Code

    Model
    Claude Opus 5.5
    Reasoning
    High
    Client version
    2.1.285 (Claude Code)
    Score
    95 · partial credit 99
    Passed
    19 of 20
    Total time
    1 h 48 min
    Total cost
    US$24.25
    Cost per pass
    US$1.28

    Subscription, at API list price.

    Claude Code in headless mode, high reasoning effort, on a Claude subscription.

  • Sonnet 5.5

    with Claude Code

    Model
    Claude Sonnet 5.5
    Reasoning
    High
    Client version
    2.1.285 (Claude Code)
    Score
    75 · partial credit 99
    Passed
    15 of 20
    Total time
    1 h 36 min
    Total cost
    US$16.63
    Cost per pass
    US$1.11

    Subscription, at API list price.

    Claude Code in headless mode, high reasoning effort, on a Claude subscription.

Results file

Downloads

History

Runs