Most AI safety work tries to make a model want the right things. Google DeepMind's latest proposal starts from a more uncomfortable premise: assume that effort will sometimes fail, and build the walls anyway. In a paper published on June 18, an AI Control Roadmap, researchers Rohin Shah and Four Flynn argue that aligning a model's goals is no longer enough on its own, and that labs need a second line of defence for the moment alignment slips.
Shah, who leads DeepMind's safety and alignment team, put it as a plain operational question. "If the first line of defence, alignment, fails, how can we mitigate harm anyway?" The answer the roadmap reaches for comes straight out of corporate security. It treats a capable, misaligned AI agent as an insider threat, the trusted employee who turns out to be working against the company.
To make that concrete, the paper borrows the MITRE ATT&CK framework, the catalogue of attacker behaviours that security teams already use to model human intruders. Applied to an AI agent, it maps out what a rogue system might attempt: exfiltrating its own model weights, tampering with logs, or quietly sabotaging the very safety research meant to keep it in check.
Defence in depth
The organising idea is defence in depth. Rather than trusting any single safeguard, DeepMind layers several on top of one another. Sandboxing keeps an agent boxed off from systems it should not touch. Endpoint monitoring watches what it actually does. Prompt-injection resistance hardens it against being hijacked through its own inputs. Detection systems sit on top, tuned to flag an agent stepping outside its lane. The point is that no one wall has to be perfect, because the next one is behind it.
DeepMind says it tested a prototype of this approach across roughly one million coding-agent tasks, the kind of repetitive, autonomous work that increasingly runs without a person checking each step. That detail is the quiet centre of the whole document. The reason to plan for misbehaving agents now is that agents are already being handed real work, at volume, with looser supervision than a human employee would ever get.
There is something bracing about a frontier lab writing down, in this much detail, the specific ways its own software might turn on it. It reads less like a manifesto and more like a threat model, which is probably the right register. The roadmap does not claim today's systems are plotting anything. It argues that the infrastructure to catch one that did should exist before the capability does, not after.
The approach fits a wider shift in how the serious labs talk about risk. Anthropic has argued for the ability to pause frontier development if systems start improving themselves too quickly. DeepMind's contribution is more mundane and, in a way, more reassuring: not a grand pause, but plumbing. Whether the rest of the industry adopts anything like it is the open question. Containment only works if everyone building powerful agents bothers to build the walls.
Sources
- i. deepmind.google
- ii. www.techtimes.com
- iii. www.axios.com
- iv. fortune.com
Commentarii · 0