Somewhere in the marketing this year, the AI agent stopped being a tool you supervise and became a colleague you delegate to. Hand it a goal, the pitch goes, walk away, come back to finished work. It is a lovely story. It is also, for anything that matters, still not true. Today's agents can take real actions in the world, and they will happily take the wrong ones without anybody watching.
Start with what an agent actually is. Strip the branding and it is a language model wired into a loop: read the situation, pick an action, run it, then read the result and pick again. That loop is genuinely useful, and it is a real step up from a chatbot that only talks, a distinction we drew when we argued an agent is more than a chatbot renamed. The trouble is what happens when the loop runs long.
Errors compound
An agent that is 95 percent reliable on a single step is not 95 percent reliable across twenty of them. Small mistakes feed the next decision, which feeds the one after, and the errors multiply instead of cancelling out. Anyone who has watched an agent confidently book the wrong meeting and then send a polite confirmation email about it has seen the compounding problem in miniature. The model is not lying. It built its third move on a flawed reading of its first, and nothing in the loop stopped it.
This is why the people actually shipping agents keep a human in the loop by design. Approval gates, spending caps and confirmation steps are not training wheels the vendors are embarrassed about. They are the thing that makes the product safe to sell. It is also worth remembering that raw model quality does not move in a straight line here. We noted recently that some newer Claude models regressed on certain tool calls, exactly the skill an autonomous agent leans on hardest.
Regulators are betting against full autonomy too
The clearest tell is that the rulebooks now assume supervision. China's implementation rules for AI agents, enforceable from July 15, set up a three-tier framework for how much authority an agent may exercise, with mandatory filings for high-risk uses. You do not build a tiered permission system for something you expect to run unattended. The regulation is written for the reality, not the pitch.
None of this means autonomy is a fantasy forever. Agents are getting better at longer tasks, and the supervision they need will keep shrinking. But shrinking is the honest word, not gone. If a product tells you its agent needs no oversight today, the safest assumption is that you are the oversight and nobody mentioned it. This is the calmer cousin of the louder fear that AI might slip human control entirely. The near-term problem is not an agent that escapes. It is one that fails quietly while you assumed it had things handled.
Sources
- i. www.upwork.com
- ii. arxiv.org
- iii. aigovernance.com
- iv. 365datascience.com
Commentarii · 0