What happened
- In late July 2026, Mercor and Ramp published APEX-Accounting on arXiv, a benchmark of 160 tasks spread across 10 simulated companies with full accounting systems.
- The tasks mirror a month-end close: reconciling accounts, accruing expenses, posting transactions and building reports. Expert accountants wrote and graded every task.
- The best result belongs to Claude Fable 5 (Max): 56.4% of the binary criteria written by experts, averaged over three of eight randomly drawn runs. Muse Spark 1.1 follows with 52.6%.
- Consistency is worse. The paper’s summary reports a best Pass@8 rate of 21.5% and says no model exceeds 2.6% of tasks solved when it must get them right in all eight runs.
Why it matters
- 56% of the criteria is not “almost right.” In a reconciliation, every unmet criterion is a difference that shows up at close, in the tax payment or in the conversation with the bank. For a small business, whoever didn’t catch that mismatch pays for it.
- The most useful part of the study is consistency. A job that comes out different on every run forces someone to compare the result with reality, and that review is a cost to budget before buying the license.
- Three rules of thumb help decide where to use an accounting agent today:
- Start with low-consequence tasks, such as sorting expenses for a first pass, and leave the final reconciliation to a person.
- Ask the vendor for traceability: which entry it made, from which document and why.
- Measure it in your own company. Take a month that is already closed, give it to the agent and compare it with the real close before using it on the current month.
- The study has limits. The companies are simulated and American, with no Chilean tax rules and no link to the SII, Chile’s tax authority. Ramp also sells financial software with AI, so it has an interest in the topic, although the paper publishes its methodology and per-model results.
The number
56.4% versus 2.6%. The best mean score per criterion, against the maximum share of tasks a model solves in all eight runs.
Context
It is the third recent measurement that separates doing from doing well. In EmailBench, email agents executed 99.7% of their calls but solved 33.5% of the tasks. Another measurement showed agents declaring themselves done 38 points more often than the evaluator passes them. The APEX-Accounting authors add a cost finding: a larger token budget raises the score, but within a single budget, tasks that consume more tokens do worse.
What’s next
- No date announced for new runs or for a version with rules outside the United States.
- Mercor keeps a public leaderboard for the benchmark, where new models would show up.
Bottom line
An accounting agent can save a small business hours. Before signing the close, its result has to sit on the same table as the accountant’s.
This note describes a published study and is not accounting or tax advice.
Sources
- APEX-Accounting, arXiv, version 2 of July 30, 2026
- Introducing APEX-Accounting, built with Ramp, Mercor
- APEX–Accounting, Ramp Labs
Written by Mamífero. Edited by Rodrigo Cornejo. See how we select and verify every note.


