This is RS9's own benchmark: the tests we ran and the results we got. It is not a research paper. We only need it to be fair, consistent and openly stated, so you can see how each score was produced and rerun the tasks yourself.
How a run works
- Same start for everyone. Each agent starts clean with the same brief, the same files and the same time limit.
- One attempt per task, up to 30 minutes each. Every attempt runs in a fresh container. The agent never sees the tests or the answers.
- The time limit is in the prompt. From the next run on, each agent is told its time limit in its instructions, so it can plan its work around it. In this first run they were not told.
- Real work, made safe. Each task is a reconstructed version of work RS9 does, built from made-up data and small repositories, so no client material is used. Research tasks use a supplied pack of documents rather than browsing the web.
- Everything is written down. For every run we record the agent's client version and settings on the day. If an agent can't be run unattended, it is left out of that run and we say so.
How it is scored
- Pass or fail against a checklist. Each task has a written list of what must be right. An attempt passes only if every item does. Items passed still count towards partial credit, explained below.
- Code and data tasks are checked by automated tests.
- Writing and research tasks are checked by one AI judge working through the checklist. The judge comes from a provider that is not one of the agents being tested or, failing that, is never the same provider as the agent it is judging. It does not know which agent wrote what.
- CAD tasks are checked by building the model the agent wrote and measuring it: the overall dimensions, the holes and cutouts, and whether the parts fit together as the spec says.
- Front-end design tasks are rendered in a real browser at desktop and phone width and measured. A vision-capable AI judge then compares the screenshots against a checklist, without knowing which agent made them.
- Human spot-checks. Dale checks a handful of judge decisions and anything that looks wrong.
- Also recorded: the time each attempt took, and its cost where there is one. An attempt that runs out of time is shown as “Timed out” and counts as not passed. An “Error” means the agent itself failed to run properly.
How we keep it fair
Before publishing, every failure in this run was reviewed against the task's written brief. Where a check demanded more than the brief actually said, we fixed the check and re-graded the saved work. No agent was run again to improve its score. A second reviewer then checked that none of those fixes had bent toward the agents' answers.
How the score works
- The headline score is out of 100. It is the percentage of the tasks an agent passed outright: 15 of 20 tasks is a score of 75. A task only counts if every item on its checklist is right.
- Partial credit tells close results apart. It gives every task equal weight and credits it with the share of its individual checks that passed, including the judge's items on writing and research tasks. Two agents on the same score can differ here, and it breaks ties in the ranking.
- Infrastructure errors are left out. If an attempt failed because our set-up or the provider broke, not because the agent got it wrong, it is not counted for or against the agent, so a score can be out of fewer than 20.
- A few points either way means little. Each task is attempted once, so one task going the other way moves a score by about eight points.
How cost is calculated
- We count tokens, then price them. For every attempt we record the tokens the agent used: what it read, what it reused from its cache, what it wrote and how much it reasoned. We multiply those by the provider's public API price list, in US dollars. Cached reading is charged at the cheaper cached rate, and reasoning at the output rate.
- Subscription agents don't bill per task. Their cost here is what the same tokens would cost at the provider's public API price, not what anyone paid. It is a fair way to compare, not a receipt.
- OpenRouter models show a real charge. For the models we run through OpenRouter, the cost is what we were actually billed. That can be well below list price, because the host serving the model may charge a discounted rate. We keep the list-price figure too, and note it on each agent's card.
- Some costs are lower bounds. When an agent is stopped at the time limit it may not report its final usage, so its cost is slightly undercounted. If an attempt reports no usage at all, that agent's total is shown as “price unknown” rather than a partial figure.
- Cost per pass is the total cost divided by the tasks passed, so an agent that is cheap but fails often does not look better than it is. If we can't find a public price, the page says “price unknown”, never $0. An agent that passes nothing has no cost per pass.
How we track changes over time
- The same tasks, rerun on a schedule. The benchmark is rerun regularly, and each run keeps its own dated page. The “Over time” chart on the results page plots every agent's tasks passed, and its cost per pass, across those runs.
- Version changes are marked. We record each agent's model and client version every run. When either differs from that agent's previous run, the point gets a diamond and the tooltip says “version changed”, so a jump can be tied to an update rather than luck.
- Read small moves with care. With one attempt per task, an agent moving up or down by a task or two between runs is normal wobble. A steady drift, or a jump right after a version change, is what is worth noticing.
Caveats
- One attempt per task. Agents are not perfectly repeatable, so a difference of a task or two between agents means very little.
- These are RS9's tasks. They are not a general measure of which agent is best, and an agent that does well here may not do well on your work.
- A dated snapshot. Each result describes the versions and settings tested that day. Agents change often, so older runs are kept as dated pages rather than updated.
- Different set-ups. Most agents are run through their own apps on a subscription. One uses a model through OpenRouter. Compare times and costs with that in mind.
- Times are not like for like. Each time includes the provider's server speed and queueing, so a slow result can reflect a busy server rather than a slow model.
Run it yourself
Each run page has a download of its full results (a JSON file), so you can check our numbers. See the latest results.
The 20 tasks
Eight categories, two or three tasks in each. Four are marked Frontier: harder tasks added after the top model passed the first set. Each task has a written brief, its starting files and a checklist of what must be right.
Software
Move a billing codebase from floats to integer minor unitsFrontier
A 125-file invoicing package must switch all money from floats to integer minor units across domain code, SQLite storage with a data migration, the JSON API, CSV import and export and reports, following a finance policy, ADRs and an issue tracker export.
Checked by: automated tests
Fix wrong totals in a checkout service
A Python e-commerce checkout has many interacting bugs behind a support ticket about wrong totals: GST rounding, promo dates across Melbourne daylight saving, a stale price cache, refunds, and the places that repeat the same sums. Fix them to the finance rules without breaking the existing tests.
Checked by: automated tests
Add multi-location stock with a safe data migration
Add multi-location stock to a live SQLite-backed stock service: migrate the production data exactly and idempotently, keep the old API byte-for-byte compatible, and add locations, transfers and per-location reports.
Checked by: automated tests
Research and decisions
Audit a vendor's marketing claims against a pack of specs, tests and standards
Label 20 marketing claims for a home battery as supported, contradicted or not determinable, naming the governing source and key figure, using 49 documents that include datasheets across revisions, lab reports, changelogs, a standard, a certification register and forum noise.
Checked by: automated tests
Six clients, one grant: eligibility and maximum funding across amendments
Work out eligibility and the exact maximum grant for six clients under a state grant whose rules were changed by three amendments, a correction and two ministerial notices, from a pack of 39 documents that includes conflicting FAQs and outdated summaries.
Checked by: AI judge against a checklist
Documents
Apply 14 agreed changes across a contract set
Fourteen negotiated changes (renumbered clauses and schedules, compounding price changes, moved dates) must be applied across an MSA, SOW, three Schedules and a Price Book, with every cross-reference and total kept right.
Checked by: automated tests
Answer a council tender from the supplier's own documents
Respond to a 30-requirement council RFT with a compliance matrix and word-limited sections, using only facts that RS9's own documents support, including a priced five-year offer worked from several sources.
Checked by: AI judge against a checklist
Price a fabrication and installation tender from a messy packFrontier
Take off quantities from DXF drawings, cost them from supplier quotes and RS9's rate card and pricing policy, and deliver an exact filled schedule of rates with GST plus a short pricing memo.
Checked by: AI judge against a checklist
Data
Analyse a checkout A/B test from raw event logs
From a raw event log, clean and analyse an A/B test of a new checkout (duplicates, bots skewing the split, device mix, a week-1 novelty effect) and give exact figures, the SRM p-value, a decision, and an honest write-up.
Checked by: AI judge against a checklist
October month-end: bank, ledger and payment processor
Reconcile a bank export, an accounting-system export and a payment processor's payouts across month-end (UTC vs Melbourne time with daylight saving), and report exact figures and every unreconciled item with its reason.
Checked by: automated tests
Interfaces and technical work
Build a three-page hardware shop from mockups and a behaviour specFrontier
Build a filterable catalogue with URL state and history, a product page and a localStorage cart with pricing rules, in plain HTML/CSS/JS, matching desktop and mobile mockups and passing hidden browser behaviour tests.
Checked by: AI judge against a checklist
Bill of materials and optimal cut list from a revised drawing set
Read text-based fabrication drawings with revisions, a draft, two change notes and a purchase order, produce a bill of materials, then an exact minimum-new-bar cut list with kerf that uses an offcut register.
Checked by: automated tests
Build a sheet-metal quote wizard to a 30-rule spec
Build a plain HTML/CSS/JS multi-step quote wizard with conditional steps, unit conversion, quantity tiers, GST and accessible errors, with the rules in a tested ES module.
Checked by: automated tests
Picking up where someone left off
Finish a feature after a five-day chat thread of changing requirements
Take over a half-built trade quote calculator, work out the currently agreed rules from a 350-message chat export full of reversals, a side-thread decision and an unconfirmed idea, finish it, and write a handoff citing the deciding messages.
Checked by: AI judge against a checklist
Diagnose a two-cause checkout incident across three services
From about 14,000 lines of logs in three time formats, metrics, deploy history and config diffs, find the two changes that combined into a retry storm (and dismiss the deploy that only looks guilty), name the first failure to the second, and propose a config fix that survives a simulated replay.
Checked by: automated tests
CAD and fabrication
Parametric sensor-board enclosure in OpenSCAD
Model a two-part electronics enclosure (base and lid) in OpenSCAD from a board drawing and a spec, with connector cutouts, standoffs, a fitted lid and fully parametric dimensions.
Checked by: automated tests
Laser-cut flat pattern for a formed bracket
Turn a formed steel bracket (dimensions and drawing) into a DXF flat pattern with bend lines and a bend table, using the given thickness, bend radius and K-factor.
Checked by: automated tests
Hinged sensor clamp and wall bracket, measured from a vendor STLFrontier
Design a three-part OpenSCAD assembly (wall bracket, hinged clamp, hinge pin) around a vendor sensor STL that must be measured, with mating clearances, aligned bolt holes, a screw stack-up, wall and stiffness minimums, a print bed limit and no interference through a range of hinge rotation.
Checked by: automated tests
Front-end design
Design a landing page that borrows an inspiration image's visual language
Build an RS9 service landing page from supplied content and fonts that clearly reuses the palette, type pairing, hard-edged components and layout motifs of an inspiration screenshot, without copying its content.
Checked by: AI judge against a checklist
Rebuild a landing page from desktop, tablet and mobile mockups
Recreate a marketing landing page from four PNG mockups (desktop, tablet, mobile, open mobile menu) with supplied brand tokens and fonts: responsive, accessible, and measured against the mockups.
Checked by: AI judge against a checklist