The month-end close is the kind of work accountants do without applause: pull the right figures out of a company's files, run the numbers, and hand back a clean table. A new study from the AI evaluation firm Mercor set that task to twelve licensed CPAs, each with about five and a half years of experience, and then set the same task to the leading AI models. On the simplified scenarios, the humans averaged roughly 37 percent. The best models scored at or near perfect.

That gap did not exist eighteen months ago. Mercor reports that OpenAI's o3 first passed the average accountant in the spring of 2025, and that GPT-5 reached about 69 percent later that year. The current frontier has pulled clear of the human baseline on these narrow exercises, and it has done so cheaply. Mercor puts the cost of a correct answer from Claude Opus 5 at around 21 cents per rubric criterion met, against roughly ten dollars for an unaided accountant doing the same work.

Where the models stop

The headline numbers come with a large asterisk. The near-perfect scores were on a trimmed set of tasks. On the fuller APEX accounting benchmark, which runs to 160 tasks across ten simulated companies, the picture is more sober. Claude Opus 5.5 led at 61.8 percent, with Fable 5.1 just behind at 61.0 percent and GPT-6 Astra at 57.9 percent. No model fully solved close to 60 percent of the tasks. As the researchers put it, AI cannot close the books without oversight yet.

There is also the question of what the benchmark measures. These tasks reward exactly what current models are good at: finding a figure buried in a spreadsheet and following an instruction to the letter. They leave out most of what fills an accountant's day. Talking a nervous client through a bad quarter, checking an odd entry with a colleague, and knowing from years of context that a number looks wrong before you can say why. None of that shows up in a rubric.

A familiar pattern

The accounting result rhymes with other professional benchmarks landing this autumn. In a separate trial reported here, AI tutors matched human experts on measured learning gains at a fraction of the cost. The shape is consistent. On the well-defined, measurable core of a knowledge job, the models are now competitive or better. On the judgment and the relationships wrapped around that core, they are not being tested, because those things are hard to score.

For firms, the practical reading is narrow and useful. The close can be drafted by a model in minutes and checked by a person, rather than built by a person from scratch. That changes how the hours are spent. It does not, on this evidence, remove the person who signs off.

Sources

  1. i. the-decoder.com
  2. ii. www.mercor.com
  3. iii. www.cfo.com

Commentarii · 0

Add · a · Comment