The standard solution to AI risk is testing. Before a frontier model is released, outside evaluators are granted some level of access and asked to probe it for harm: hazardous biological knowledge, cyber capability, deceptive behavior, the full menu of capability concerns. The principle is sound. The implementation, a new report argues, is becoming a security liability of its own.

The Royal United Services Institute, the venerable UK defense think tank, published its analysis of the third-party evaluation ecosystem this month. Its core finding is that nobody quite agrees what "secure access" means. One evaluator may receive a tightly scoped API and a non-disclosure agreement. Another may get deeper visibility into model internals, training environments, or production infrastructure. Standards differ between labs, between governments, and between commercial auditors.

That looseness creates exploitable seams. The Register summarised the report's central concern starkly: the act of giving outsiders deep visibility into the most capable models on Earth is itself a new attack surface. Hostile states and organised criminal groups have obvious incentives to slip into evaluator privileges if the bar is low enough, and insider risk compounds the problem.

The Access-Risk Matrix

RUSI's proposed remedy is what it calls an Access-Risk Matrix that maps types of access against threat models. At one end sits read-only API testing of a model's text outputs, which carries modest risk. At the other end sits write access to model internals, which the report flags as the highest-risk category. An adversary in that position could alter a model's behavior directly, with consequences that would propagate downstream to every user who later calls the model.

The report does not suggest that evaluations should stop. It argues, rather, that the system needs the kind of formalized governance long established for sensitive government work: vetted personnel, classified evaluation environments, and explicit cross-jurisdictional standards for what kind of access produces what kind of accountability.

The mathematics of testing has shifted

This is a useful counterweight to the assumption that more testing automatically produces more safety. If the testing infrastructure itself becomes a soft target, the security mathematics inverts. The marginal evaluator with privileged access is no longer purely a safety gain. They are also a potential point of compromise, and the bigger and looser the evaluator pool gets, the worse the worst-case risk becomes.

The wider question is whether the major labs and their evaluators will treat the RUSI matrix as a serious starting point or as another white paper to file away. Five Eyes intelligence agencies recently told organisations to treat AI agents as unsafe by default. The same caution may need to apply to the people who test them. Whether the industry can rebuild evaluator access from the ground up, while the deployment race continues at its current pace, is the question this report leaves on the table.

Sources

  1. i. www.theregister.com
  2. ii. www.rusi.org
  3. iii. www.rusi.org
  4. iv. www.axios.com
  5. v. www.paloaltonetworks.com

Commentarii · 0

Add · a · Comment