# AI vs Accountants: Mercor's CPA Study, Read Closely

> Mercor's study put 12 CPAs against Claude Opus 5: 37% vs 100% on four close tasks. What it measured, what it did not, and the full benchmark scores.

- Author: Nitish Garg, Founder & CEO, CellCog
- Published: 2026-10-03
- Canonical (HTML): https://cellcog.ai/blog/ai-vs-accountants-mercor-study/
- Section: Guides / Workflows & use cases
- Publisher: CellCog (https://cellcog.ai), the AI employee platform. Blog index for agents: https://cellcog.ai/blog/llms.txt

## Key points

- Mercor published a study on October 1, 2026 in which 12 licensed US CPAs, averaging 5.4 years of accounting experience, completed four month-end close tasks adapted from its APEX-Accounting benchmark.
- Working without AI, the accountants averaged about 37% of rubric criteria; Claude Opus 5 working alone scored 100% on all 20 attempts and finished each in under 10 minutes.
- Per rubric criterion met, Mercor puts Opus 5 at $0.21 against $10.35 for an accountant at the US median wage, 49 times more; the model figure is a list-price upper bound.
- Accountants working with Claude were slower and slightly less accurate than Claude alone, taking about 15 times as long per task.
- The tasks were simplified versions of benchmark tasks built to trip up models, with no colleagues and no client to ask; Mercor and its task authors say the setting and scoring drove the gap.
- On the full APEX-Accounting benchmark, 160 tasks across 10 company worlds, the same Opus 5 meets 54.5% of criteria and the best model listed, Opus 5.5 at Max effort, 61.8%.

## At a glance

- **What did the Mercor study find?** On four simplified month-end close tasks, 12 licensed CPAs averaged about 37% of rubric criteria without AI, while Claude Opus 5 alone scored 100% on 20 attempts in under 10 minutes each.
- **Does it mean AI can replace accountants?** Mercor says no. The tasks tested detail work, file search and instruction following, and stripped out colleagues, client questions and job context, which is much of the real job.
- **How much cheaper was the AI?** Mercor estimates $0.21 per rubric criterion met for Opus 5 against $10.35 for an accountant at the US median wage, a 49 times gap, with the model cost an upper bound.
- **How do models do on the full benchmark?** Much lower. On APEX-Accounting the leaderboard lists Opus 5.5 at Max effort at 61.8% and Opus 5 at 54.5%, and Mercor says 58% of tasks are never fully solved.

Twelve licensed CPAs against one AI model, on the same four accounting tasks, graded on the same rubric. Mercor published the result on October 1, 2026: the accountants averaged about 37% of the criteria, and Claude Opus 5 scored 100% on every one of its 20 attempts. The [blog post](https://www.mercor.com/blog/human-baselines-for-benchmarks-ai-now-outperforms-junior-accountants/) concludes that "frontier AI models are now faster and more accurate than junior accountants", and the [author's thread on X](https://x.com/aden_barton/status/2105766679481553157) passed 200,000 views within a day. Mercor itself flagged the risk of reading too much into it: "So much so that we considered not publishing them for fear of misinterpretation." Here is what the [15-page paper](https://cdn.sanity.io/files/h6s14f4z/production/4048fb3eece18a56006026d212d25312f7ffc29a.pdf) actually measured, what it left out, and how the same models score on the full benchmark the tasks came from.

## What Mercor tested

*Table: The study design, per Mercor's paper of October 2026*

| Element | What Mercor did |
|---|---|
| Participants | 12 US accountants, all with an active CPA license; 5.4 years of accounting experience on average (range 3.8 to 8.0) |
| Seniority | 2 staff or associate, 7 senior, 3 manager or above; no partners or directors |
| Tasks | 4 month-end close scenarios adapted and simplified from APEX-Accounting, each on a simulated company's working files |
| Time and pay | A 3-hour window per task, paid for the full 3 hours, plus a bonus per criterion met |
| AI help | Each accountant did 2 tasks alone and 2 with Claude Cowork running Opus 5, in random order |
| AI alone | Models ran as tool-using agents capped at 250 steps, on the same files, rubrics and grader prompts |
| Grading | Rubrics of 13 to 22 items per task written by accountants, scored by an AI judge; 98% agreement with author hand-grades on 8 submissions |

The four scenarios were a nonprofit's budget-versus-actual with a shared cost pool to allocate, a hotel whose booking system and books disagreed on December room revenue, a venue's event revenue variance, and a one-month roll-forward of three lease schedules. One hotel rubric item, for example, required spotting $4,000 of parking revenue miscoded as room revenue.

## The results

*Table: Accuracy, time and cost by group, per Mercor's paper*

| Group | Task attempts | Accuracy (share of rubric criteria met) | Time per task | Cost per criterion met |
|---|---|---|---|---|
| Accountants without AI | 23 | About 37% on average; spread from 0% to about 90% | Mostly 30 to 180 minutes | $10.35 at the US median accountant wage |
| Accountants with Claude | 24 | Slightly below Claude alone | About 15 times Claude alone | Not published |
| Claude Opus 5 alone | 20 | 100% on every attempt | Under 10 minutes each | $0.21, list prices, no caching discount |

Even the best accountant in the sample scored below the model, and most participants submitted before their time ran out, so Mercor reads time as not the main constraint. The 49 times cost gap compares a model priced at list rates against the Bureau of Labor Statistics median accountant wage, not the rate the participants were paid. On the narrower set of criteria humans usually got right, those with a pass rate above 60%, Mercor reports a 34 times speedup for the model working alone.

The timeline matters more to Mercor than the gap. Its chart of every model it tested by release date puts GPT-4o near 0%, OpenAI's o3 above the 37% accountant line in spring 2025, GPT-5 at about 69%, and Opus 5 and other recent frontier models at or near 100%. Budget and open-weight models such as Qwen3.5-122B still sit below the accountant line. The author's own summary on X: "As of May 2025, accountants outperformed AI."

## The catch: these were simplified tasks

The study tasks were pared-down versions of APEX-Accounting tasks, and the file map Mercor handed participants cut search time further. On the full benchmark the numbers are very different. Mercor's [APEX-Accounting leaderboard](https://www.mercor.com/apex/apex-accounting-leaderboard/) covers 10 company worlds, 160 tasks and 2,186 criteria, and says 58% of tasks are never fully solved by any model.

*Table: Top of the APEX-Accounting leaderboard, mean score, read October 3, 2026*

| Model and effort | Mean score |
|---|---|
| Claude Opus 5.5, Max | 61.8% (plus or minus 4.0) |
| Claude Fable 5.1, Max | 61.0% (plus or minus 3.7) |
| Claude Opus 5.5, Medium | 59.9% (plus or minus 3.7) |
| GPT-6 Astra, Max | 57.9% (plus or minus 3.8) |
| Claude Fable 5, Max | 56.4% (plus or minus 3.6) |
| Claude Opus 5, Max | 54.5% (plus or minus 3.9) |

The same model that aced the simplified study meets about half the criteria on the full benchmark, where Mercor's own takeaway is that frontier agents "cannot yet reliably close the books". The leaderboard page also carries an inconsistency worth noting: its header shows a highest score of 61.0% while the table lists Opus 5.5 at 61.8%, and its takeaways still name Fable 5 as the best model. We quote the table rows.

## Why the CPAs lost, per Mercor

Mercor went back to its task authors after the results and concluded the tasks were fair: the right answer was findable in the files. The gap, it argues, came from three things.

- **The tasks were built to fail models.** They stacked hard-to-spot but realistic traps, like hotel stays billed on January 3 that belonged in December revenue, or a monthly lease reclassification step most participants skipped. One missed number could cascade into a very low score. The task authors expected juniors to score about 30% and mid-level accountants about 55%.
- **The setting stripped out the job.** No colleagues, no supervisor, no client to ask, no months of context. One task author put it this way: "what is unrealistic is the setting and the scoring".
- **The tasks tested what models do best.** Detail work, searching every file and following instructions closely. One author observed that people "tended to settle early on an approach that looked reasonable" and carried it through.

The participants said the same in their own words. Seven of the 12 rated the tasks realistic; the two who rated them very unrealistic later clarified their objection was the setting, not the spreadsheet work. One wrote: "In reality I can usually ask my client questions as I work".

## AI plus an accountant did worse than AI alone

The study was designed to measure how much Claude speeds up an accountant. It could not: because Claude alone scored 100%, the rubric had no room to reward human judgment, and in Mercor's words "humans with AI are slower and slightly less accurate than AI alone". The paper's most useful detail is in the misses. In three of the four assisted sessions that fell short of a perfect score, Claude had the correct figure at some point; twice Claude revised away from it and the accountant deferred, and once the accountant overrode Claude's correct first answer. Review is a skill, and on a task the model already gets right, a reviewer adds time and risk.

## What it means for a small business's books

Mercor's own reading is that accounting work will move toward client relationships, judgment under uncertainty and asking the right questions, while the detail-heavy close work gets automated. For a small business, the practical version is simpler: the reconciliations, variance tables and journal entries in these four tasks are the work an AI can already do fast and carefully, and the parts that need context about your business are the parts you still own. CellCog runs its AI employees on Claude Opus 5.5, the top model on the APEX-Accounting table above, and an [AI bookkeeper](https://cellcog.ai/ai-employees/ai-bookkeeper) works the way the study's best setup did: the model does the close work, and you review against your own context. Our breakdown of [what an AI employee costs](https://cellcog.ai/blog/ai-employee-cost/) covers how that is priced.

## What we are watching for

- A follow-up study with harder, less simplified tasks, or with colleagues and client questions allowed.
- Human baselines on the full APEX-Accounting benchmark, not only the simplified tasks.
- Whether Mercor publishes the assisted group's exact accuracy and the per-task scores.
- New leaderboard entries above 61.8%, and whether the 58% never-solved share falls.
- Similar human-baseline studies from Mercor's other APEX benchmarks.

## The record

As of October 3, 2026, 02:00 UTC: page opened. We read Mercor's blog post (published October 1, 2026 at 18:59 UTC per its page metadata), the full paper PDF linked from it, and the APEX-Accounting leaderboard directly on October 3 between 01:42 and 01:47 UTC. Aden Barton's thread was read on X the same hour; its first post is timestamped October 1, 2026 at 21:07 UTC. Model positions on Mercor's release-date chart are taken from Mercor's own figure text, not measured by us.

## FAQ

**Who ran the study and when was it published?**

Mercor, the AI training-data company, published it on October 1, 2026. The paper is by Aden Barton, a Mercor researcher, and the tasks come from Mercor's APEX-Accounting benchmark.

**Were the accountants qualified?**

All 12 held an active US CPA license and an accounting degree, 8 held a master's degree, they averaged 5.4 years of experience, and half had worked at a Big Four firm. None were partners or directors.

**How were the answers graded?**

Against rubrics of 13 to 22 items per task written by accountants, scored by an AI judge. Mercor says eight submissions hand-graded by the task authors agreed with the judge 98% of the time.

**Why did accountants using Claude score lower than Claude alone?**

Because Claude alone already scored 100%, the rubric could not reward any human improvement, and review time only slowed things down. In three of the four imperfect assisted sessions, Claude had the right figure at some point but it did not reach the final answer.

**Which tasks were in the study?**

Four month-end close scenarios: a nonprofit's budget-versus-actual with shared cost allocation, a hotel's room-revenue and commission reconciliation, a venue's event revenue variance, and a roll-forward of three lease schedules under US lease accounting.

## Related

- [What Is an AI Employee? The 5-Part Test for a Standing AI Worker](https://cellcog.ai/blog/what-is-an-ai-employee/index.md)
- [Anthropic's Three Numbers: Claude Leads 26% of Its AI R&D](https://cellcog.ai/blog/anthropic-rd-automation-index/index.md)
- [How Much Does an AI Employee Cost? A 7-Layer Total-Cost Framework](https://cellcog.ai/blog/ai-employee-cost/index.md)

## The AI employee for this read

[AI Bookkeeper](https://cellcog.ai/ai-employees/ai-bookkeeper): Categorizes, reconciles and flags what does not add up, every day.

---

Markdown alternate of https://cellcog.ai/blog/ai-vs-accountants-mercor-study/. Try CellCog free, no credit card needed: https://cellcog.ai/signup
