The demos are genuinely good. An agent reads a ticket, checks three systems, drafts a reply and files the result, all on its own. Then a company tries to run that same agent across a real workflow a few thousand times, and the wheels come off. This gap between the demo and the deployment has become the defining enterprise story of 2026, and the numbers behind it are starting to firm up.

The pattern in this year's benchmarking is consistent. Leading models handle single step tasks well, often scoring in the eighties or nineties. Ask them to sustain a workflow across several applications and many steps, and success rates fall sharply, in some tests to the rough range of a fifth to a quarter of attempts. The cause is not mystery so much as arithmetic. Chain enough steps together and small error rates compound. An agent that is 90 percent reliable at each step of a ten step job succeeds end to end only about a third of the time. Drop to 85 percent and you are down near one in five.

The evaluation gap

Part of the trouble is that teams cannot see the failures until they hit production. Writing in VentureBeat, analysts describe a reality alignment problem: most organisations are not short on test coverage, they are short on tests that resemble the messy conditions agents actually meet. And a striking share ship anyway. Survey work compiled by Writer found the majority of enterprises reporting real difficulty turning agent pilots into dependable systems, even as budgets kept climbing.

None of this means agents do not work. It means the honest unit of measurement is the whole task, not the single step, and that the last stretch of reliability is the expensive one. The companies making progress tend to narrow the job, keep a human on the tricky decisions, and instrument every run so failures surface early rather than in front of a customer.

Why it matters

The stakes are practical. Agents that act on your behalf raise a plain question about who is actually in control when something goes wrong, which is one reason courts and regulators have started weighing in on what an agent is allowed to do. It is also why a market has sprung up for watching and testing agents before they touch anything important. The technology is real. The reliability is a build, not a given.

Sources

  1. i. venturebeat.com
  2. ii. writer.com
  3. iii. temporal.io

Commentarii · 0

Add · a · Comment