Skip to content

Can AI Replace Accountants? What Claude Opus 5 vs. 12 CPAs Really Found

Nifty Tech Finds
9 min read
14 views

Can AI replace accountants? Not yet, and the data behind the headline is more interesting than the simple yes-or-no answer. A new benchmark from the AI talent marketplace Mercor found that Anthropic’s Claude Opus 5 scored a perfect 100% on a simplified month-end accounting test, far ahead of the 37% average posted by 12 licensed CPAs. But a separate, harder version of the same benchmark — one built to mirror real, messy accounting work — shows the best AI models still fail more than half of the tasks outright. Here’s what Mercor actually tested, what the numbers mean, and why “AI beats accountants” is only half the story.

Key Takeaways

  • On a simplified Mercor benchmark, Claude Opus 5 hit 100% accuracy on all 20 attempts at month-end close tasks, versus a 37% average for 12 licensed CPAs with 5.5 years of experience.
  • Claude Opus 5 finished each task in under 10 minutes for about $0.21 per graded criterion, compared with 30–180 minutes and $10.35 per criterion for the human accountants — roughly 49 times cheaper.
  • On APEX-Accounting, Mercor’s harder, more realistic benchmark of 160 tasks across 10 simulated businesses, the top-scoring AI models manage only about 55–62% accuracy.
  • Researchers found that 58% of APEX-Accounting’s tasks are never fully solved by any model, and that judgment — not tool access or raw spending — is the main bottleneck.
  • Both studies excluded client communication, ambiguous requests, and collaborative problem-solving — the parts of accounting that don’t show up in a graded worksheet.

What Mercor Actually Tested

Mercor is a talent and research company that builds benchmarks to measure how AI models perform on real professional work, not academic trivia. In early October 2026, it published two related but distinct pieces of research on AI and accounting, and conflating them is how a lot of the week’s headlines got oversimplified.

The first is a focused study comparing Claude Opus 5 against 12 licensed, practicing CPAs on four simplified month-end close scenarios. The second is APEX-Accounting, a much larger and harder benchmark of 160 tasks spread across 10 simulated businesses, designed to test whether AI agents can handle the kind of long, multi-application accounting work a junior employee actually does. (Source)

Calculator and laptop on a desk, illustrating whether AI can replace accountants

The Simplified Test: Claude Opus 5 vs. 12 CPAs

For the first study, Mercor hired 12 licensed CPAs averaging 5.5 years of experience and had them complete four month-end close scenarios: finding figures buried in company files, running the right calculations, and delivering clean, tabular results. The tasks deliberately included hard-to-spot requirements that compounded on each other, so one missed number could tank an otherwise solid answer.

The humans scored anywhere from 0% to 90%, averaging 37%. Claude Opus 5 scored 100% on all 20 of its attempts. It also worked faster — under 10 minutes per task versus 30 to 180 minutes for the CPAs — and far cheaper, at roughly $0.21 per graded rubric criterion compared with $10.35 for a human accountant. Mercor’s own researchers note this is a sharp jump from where things stood even recently: AI models from mid-2024 scored close to 0% on similar tasks, and it wasn’t until OpenAI’s o3 model arrived in spring 2025 that any AI system cleared the human average.

The Harder Test: APEX-Accounting’s Real-World Tasks

The second study tells a less triumphant story. APEX-Accounting evaluates AI agents across 10 simulated businesses containing 160 distinct tasks and 2,186 individual grading criteria, built to test whether an AI agent can “reason over incomplete records, apply company-specific context, use multiple applications, and carry figures accurately through a close.” (Source)

On this version of the test, even the best-performing models land in the 55–62% accuracy range, not the 100% Claude Opus 5 hit on the simplified test. Mercor reports that 58% of the benchmark’s tasks are never fully solved by any model across any number of attempts, and that models are inconsistent — a model that nails a task once often fails the same task on a second try. The researchers also found that spending more on a given attempt doesn’t reliably buy better accuracy, which points to accounting judgment, not tool access or compute budget, as the real limiting factor.

Why the Two Results Don’t Contradict Each Other

A model scoring 100% on one Mercor test and 55–62% on another isn’t a contradiction once you look at what each test is actually measuring. The simplified study isolated exactly the kind of work AI is already strong at: locating numbers, following explicit instructions precisely, and doing arithmetic without fatigue or distraction — the same reason AI models have posted strong scores on other structured knowledge work. APEX-Accounting instead strings many of those smaller skills together into long, cross-application tasks that require holding context over time and filling in gaps the way an experienced accountant would, which is a much harder problem for today’s models.

Mercor’s own researchers are explicit about this limitation: their simplified benchmark tested “a set of accounting skills that happen to be the same skills models are best at,” and it excluded client communication, collaborative problem-solving, and the contextual judgment that make up a meaningful share of a real accountant’s job. Read together, the two studies say something more precise than “AI beats accountants”: AI has become extremely good at the narrow, rules-based slice of accounting work, and still struggles with the broader, judgment-heavy slice that CPAs are actually paid for.

Can AI Replace Accountants? What This Means for You

For accounting firms and finance teams, the practical takeaway isn’t that AI is coming for every accounting job at once. It’s that the structured, repetitive parts of the job — reconciling figures, drafting first-pass month-end numbers, flagging obvious errors — are now squarely within reach of a well-prompted AI model, at a fraction of the time and cost of a junior hire. That tracks with a broader pattern across white-collar work in 2026, where AI tools have moved from drafting assistance to handling entire structured tasks end to end, a shift we’ve also covered in how AI agents are already being put to work generating revenue rather than just saving time.

For individual accountants, the more useful benchmark isn’t the simplified one that made headlines — it’s APEX-Accounting, where even frontier models leave more than half of realistic tasks unsolved. That gap is exactly where human judgment, client relationships, and the ability to say “this number looks wrong, let me ask why” still earn their keep. If your firm is experimenting with AI for bookkeeping or month-end work, the Mercor data suggests it’s worth automating the narrow, well-defined pieces first and keeping a human reviewing anything that touches ambiguity, unusual transactions, or a client conversation.

This also fits into the wider story of how fast frontier models have been improving generally this quarter — see our roundup of September–October AI model and tool updates for the bigger picture, and our look at how today’s leading AI agents compare if you’re deciding which model to actually trust with real work.

Frequently Asked Questions

Did Claude Opus 5 really score 100% against licensed accountants?

Yes, on a specific, simplified Mercor benchmark of four month-end close scenarios. Claude Opus 5 scored 100% accuracy across all 20 attempts, while 12 licensed CPAs averaged 37% on the same tasks. That result does not apply to Mercor’s harder, more realistic accounting benchmark.

What is Mercor’s APEX-Accounting benchmark?

APEX-Accounting is a harder, more realistic benchmark with 160 tasks across 10 simulated businesses and 2,186 grading criteria, built to test whether AI agents can handle long, multi-application accounting work the way a junior employee would. The top AI models score only about 55–62% on it.

Can AI replace accountants?

Not based on this data. AI has become very strong at narrow, rules-based accounting tasks like reconciling figures or drafting month-end numbers, but on Mercor’s harder benchmark, 58% of realistic tasks are never fully solved by any model. Client communication and contextual judgment remain the bottleneck.

How much cheaper is AI than a human accountant, according to the study?

In Mercor’s simplified benchmark, Claude Opus 5 cost about $0.21 per graded rubric criterion and finished tasks in under 10 minutes, versus $10.35 per criterion and 30 to 180 minutes for the human CPAs — roughly 49 times more expensive for the humans on that specific test.

What parts of accounting can AI not do yet?

Mercor’s researchers note that both of its studies excluded client communication, collaborative problem-solving, and the contextual judgment calls that make up a real accountant’s job. On the harder APEX-Accounting benchmark, AI agents still fail more than half of realistic, cross-application tasks outright.

The honest summary of Mercor’s research isn’t “AI beats your accountant” or “AI can’t do accounting” — it’s that AI has gotten remarkably good at a specific, structured slice of the job while still missing most of the harder, judgment-heavy work that defines it. That pattern — huge gains on narrow benchmarks, much more modest gains on messy, real-world tasks — is showing up across white-collar AI research this year, and accounting just happens to be where it’s easiest to put a number on it.

Leave a Comment

Monthly digest

Get the month in AI, once a month

One email a month with everything worth reading from NiftyTechFinds. No spam, no daily pings, unsubscribe in one click.

LET’S KEEP IN TOUCH!

We’d love to keep you updated with our latest news and offers 😎

We don’t spam! Read our privacy policy for more info.