With the White House now testing frontier models for their cyber capabilities, it is easy to picture the threat as a machine that breaks into systems entirely on its own, no human required. It is a vivid image. It is also, for the moment, mostly fiction. The Center for Strategic and International Studies put it bluntly: fully autonomous AI cyberattacks remain largely aspirational.
That word, aspirational, is doing a lot of work, and it deserves an honest unpacking rather than a flat dismissal. The capability is not zero. In controlled testing, Anthropic's offensive-security model reportedly developed working exploits with a success rate above 70 percent. That is genuinely striking. But a high score in a lab, where a model is pointed at a problem and told to solve it, is not the same as a system that picks its own targets, adapts when things go wrong, and sees an intrusion through from start to finish without a person validating each step. CSIS is clear on the gap: human operators remain central to directing and checking what these models produce.
What the real cases actually show
The strongest evidence for AI-driven attacks tends to get flattened in the retelling. Anthropic disclosed a case in late 2025 in which an AI agent carried out tasks across an espionage operation, from reconnaissance through to data exfiltration. Real, confirmed, and significant. It was also one of the first cases of its kind rather than a sign of something widespread, and it was attributed to a well-resourced state-backed group, not a lone amateur.
A more representative episode is the breach of nine Mexican government agencies, where a single operator leaned on tools like Claude and ChatGPT. The detail that matters: the operator still directed every stage of the intrusion themselves. What the AI compressed was time, not judgement. In one instance the gap between a model refusing a request and being talked into it was about 40 minutes. That is a real problem. It is a different problem from autonomy.
The boring threat is the bigger one
Strip away the science-fiction framing and what AI actually changes is speed, scale, and access. Attackers now scan for newly disclosed vulnerabilities within minutes of a CVE going public, sometimes before defenders have finished reading the advisory. Phishing messages drafted by a model are several times more likely to get a click. And the gap between a middling attacker and an elite one narrows when a model handles the fiddly parts. Researchers measuring AI agents on multi-step attack scenarios, in work posted to arXiv, find steady progress on pieces of the kill chain rather than a single system that does the whole job.
There is even a self-correcting check built into the hype. AI-generated malware often contains errors, and the less skilled operators most likely to lean on it are the least able to spot when the model has handed them something broken. The same week these fears circulate, recall that at this year's Pwn2Own, AI coding tools were the targets that fell, not the attackers that won.
None of this is a reason to relax. The direction of travel is real, and a model that needs less and less hand-holding is a reasonable thing to plan for. But CSIS warns against writing policy around the worst-case movie plot, because doing so pulls attention away from the steady, unglamorous erosion of the defender's advantage that is already happening. The honest summary: AI is making human attackers faster and cheaper today. The fully autonomous hacker is a forecast, not a fact.
Sources
- i. www.csis.org
- ii. www.securityweek.com
- iii. arxiv.org
- iv. foresiet.com
Commentarii · 0