Twelve licensed CPAs against one AI model, on the same four accounting tasks, graded on the same rubric. Mercor published the result on October 1, 2026: the accountants averaged about 37% of the criteria, and Claude Opus 5 scored 100% on every one of its 20 attempts. The blog post concludes that “frontier AI models are now faster and more accurate than junior accountants”, and the author’s thread on X passed 200,000 views within a day. Mercor itself flagged the risk of reading too much into it: “So much so that we considered not publishing them for fear of misinterpretation.” Here is what the 15-page paper actually measured, what it left out, and how the same models score on the full benchmark the tasks came from.
On this page · 8 sectionsOpen
- Mercor published a study on October 1, 2026 in which 12 licensed US CPAs, averaging 5.4 years of accounting experience, completed four month-end close tasks adapted from its APEX-Accounting benchmark.
- Working without AI, the accountants averaged about 37% of rubric criteria; Claude Opus 5 working alone scored 100% on all 20 attempts and finished each in under 10 minutes.
- Per rubric criterion met, Mercor puts Opus 5 at $0.21 against $10.35 for an accountant at the US median wage, 49 times more; the model figure is a list-price upper bound.
- Accountants working with Claude were slower and slightly less accurate than Claude alone, taking about 15 times as long per task.
- The tasks were simplified versions of benchmark tasks built to trip up models, with no colleagues and no client to ask; Mercor and its task authors say the setting and scoring drove the gap.
- On the full APEX-Accounting benchmark, 160 tasks across 10 company worlds, the same Opus 5 meets 54.5% of criteria and the best model listed, Opus 5.5 at Max effort, 61.8%.
§ 01What Mercor tested
| Element | What Mercor did |
|---|---|
| Participants | 12 US accountants, all with an active CPA license; 5.4 years of accounting experience on average (range 3.8 to 8.0) |
| Seniority | 2 staff or associate, 7 senior, 3 manager or above; no partners or directors |
| Tasks | 4 month-end close scenarios adapted and simplified from APEX-Accounting, each on a simulated company’s working files |
| Time and pay | A 3-hour window per task, paid for the full 3 hours, plus a bonus per criterion met |
| AI help | Each accountant did 2 tasks alone and 2 with Claude Cowork running Opus 5, in random order |
| AI alone | Models ran as tool-using agents capped at 250 steps, on the same files, rubrics and grader prompts |
| Grading | Rubrics of 13 to 22 items per task written by accountants, scored by an AI judge; 98% agreement with author hand-grades on 8 submissions |
The four scenarios were a nonprofit’s budget-versus-actual with a shared cost pool to allocate, a hotel whose booking system and books disagreed on December room revenue, a venue’s event revenue variance, and a one-month roll-forward of three lease schedules. One hotel rubric item, for example, required spotting $4,000 of parking revenue miscoded as room revenue.
§ 02The results
| Group | Task attempts | Accuracy (share of rubric criteria met) | Time per task | Cost per criterion met |
|---|---|---|---|---|
| Accountants without AI | 23 | About 37% on average; spread from 0% to about 90% | Mostly 30 to 180 minutes | $10.35 at the US median accountant wage |
| Accountants with Claude | 24 | Slightly below Claude alone | About 15 times Claude alone | Not published |
| Claude Opus 5 alone | 20 | 100% on every attempt | Under 10 minutes each | $0.21, list prices, no caching discount |
Even the best accountant in the sample scored below the model, and most participants submitted before their time ran out, so Mercor reads time as not the main constraint. The 49 times cost gap compares a model priced at list rates against the Bureau of Labor Statistics median accountant wage, not the rate the participants were paid. On the narrower set of criteria humans usually got right, those with a pass rate above 60%, Mercor reports a 34 times speedup for the model working alone.
The timeline matters more to Mercor than the gap. Its chart of every model it tested by release date puts GPT-4o near 0%, OpenAI’s o3 above the 37% accountant line in spring 2025, GPT-5 at about 69%, and Opus 5 and other recent frontier models at or near 100%. Budget and open-weight models such as Qwen3.5-122B still sit below the accountant line. The author’s own summary on X: “As of May 2025, accountants outperformed AI.”
§ 03The catch: these were simplified tasks
The study tasks were pared-down versions of APEX-Accounting tasks, and the file map Mercor handed participants cut search time further. On the full benchmark the numbers are very different. Mercor’s APEX-Accounting leaderboard covers 10 company worlds, 160 tasks and 2,186 criteria, and says 58% of tasks are never fully solved by any model.
| Model and effort | Mean score |
|---|---|
| Claude Opus 5.5, Max | 61.8% (plus or minus 4.0) |
| Claude Fable 5.1, Max | 61.0% (plus or minus 3.7) |
| Claude Opus 5.5, Medium | 59.9% (plus or minus 3.7) |
| GPT-6 Astra, Max | 57.9% (plus or minus 3.8) |
| Claude Fable 5, Max | 56.4% (plus or minus 3.6) |
| Claude Opus 5, Max | 54.5% (plus or minus 3.9) |
The same model that aced the simplified study meets about half the criteria on the full benchmark, where Mercor’s own takeaway is that frontier agents “cannot yet reliably close the books”. The leaderboard page also carries an inconsistency worth noting: its header shows a highest score of 61.0% while the table lists Opus 5.5 at 61.8%, and its takeaways still name Fable 5 as the best model. We quote the table rows.
§ 04Why the CPAs lost, per Mercor
Mercor went back to its task authors after the results and concluded the tasks were fair: the right answer was findable in the files. The gap, it argues, came from three things.
- The tasks were built to fail models. They stacked hard-to-spot but realistic traps, like hotel stays billed on January 3 that belonged in December revenue, or a monthly lease reclassification step most participants skipped. One missed number could cascade into a very low score. The task authors expected juniors to score about 30% and mid-level accountants about 55%.
- The setting stripped out the job. No colleagues, no supervisor, no client to ask, no months of context. One task author put it this way: “what is unrealistic is the setting and the scoring”.
- The tasks tested what models do best. Detail work, searching every file and following instructions closely. One author observed that people “tended to settle early on an approach that looked reasonable” and carried it through.
The participants said the same in their own words. Seven of the 12 rated the tasks realistic; the two who rated them very unrealistic later clarified their objection was the setting, not the spreadsheet work. One wrote: “In reality I can usually ask my client questions as I work”.
§ 05AI plus an accountant did worse than AI alone
The study was designed to measure how much Claude speeds up an accountant. It could not: because Claude alone scored 100%, the rubric had no room to reward human judgment, and in Mercor’s words “humans with AI are slower and slightly less accurate than AI alone”. The paper’s most useful detail is in the misses. In three of the four assisted sessions that fell short of a perfect score, Claude had the correct figure at some point; twice Claude revised away from it and the accountant deferred, and once the accountant overrode Claude’s correct first answer. Review is a skill, and on a task the model already gets right, a reviewer adds time and risk.
§ 06What it means for a small business’s books
Mercor’s own reading is that accounting work will move toward client relationships, judgment under uncertainty and asking the right questions, while the detail-heavy close work gets automated. For a small business, the practical version is simpler: the reconciliations, variance tables and journal entries in these four tasks are the work an AI can already do fast and carefully, and the parts that need context about your business are the parts you still own. CellCog runs its AI employees on Claude Opus 5.5, the top model on the APEX-Accounting table above, and an AI bookkeeper works the way the study’s best setup did: the model does the close work, and you review against your own context. Our breakdown of what an AI employee costs covers how that is priced.
§ 07What we are watching for
- A follow-up study with harder, less simplified tasks, or with colleagues and client questions allowed.
- Human baselines on the full APEX-Accounting benchmark, not only the simplified tasks.
- Whether Mercor publishes the assisted group’s exact accuracy and the per-task scores.
- New leaderboard entries above 61.8%, and whether the 58% never-solved share falls.
- Similar human-baseline studies from Mercor’s other APEX benchmarks.
§ 08The record
As of October 3, 2026, 02:00 UTC: page opened. We read Mercor’s blog post (published October 1, 2026 at 18:59 UTC per its page metadata), the full paper PDF linked from it, and the APEX-Accounting leaderboard directly on October 3 between 01:42 and 01:47 UTC. Aden Barton’s thread was read on X the same hour; its first post is timestamped October 1, 2026 at 21:07 UTC. Model positions on Mercor’s release-date chart are taken from Mercor’s own figure text, not measured by us.
Q1Who ran the study and when was it published?
Mercor, the AI training-data company, published it on October 1, 2026. The paper is by Aden Barton, a Mercor researcher, and the tasks come from Mercor’s APEX-Accounting benchmark.
Q2Were the accountants qualified?
All 12 held an active US CPA license and an accounting degree, 8 held a master’s degree, they averaged 5.4 years of experience, and half had worked at a Big Four firm. None were partners or directors.
Q3How were the answers graded?
Against rubrics of 13 to 22 items per task written by accountants, scored by an AI judge. Mercor says eight submissions hand-graded by the task authors agreed with the judge 98% of the time.
Q4Why did accountants using Claude score lower than Claude alone?
Because Claude alone already scored 100%, the rubric could not reward any human improvement, and review time only slowed things down. In three of the four imperfect assisted sessions, Claude had the right figure at some point but it did not reach the final answer.
Q5Which tasks were in the study?
Four month-end close scenarios: a nonprofit’s budget-versus-actual with shared cost allocation, a hotel’s room-revenue and commission reconciliation, a venue’s event revenue variance, and a roll-forward of three lease schedules under US lease accounting.
