For the better part of a year, some of the most alarming messages arriving in OpenAI's ChatGPT, Google's Gemini and Character.AI came from accounts that looked like distressed teenagers. They were not. According to an investigation by WIRED, they were the work of contractors hired on behalf of Meta, feeding crisis prompts about suicide, self-harm, drugs and sex into rival chatbots to see how the safety systems held up.

The project ran internally under the name Cannes, and a contracting firm called Covalen handled the day-to-day work. WIRED reports that hundreds of contractors, many of them based in Kenya, set up dummy accounts registered to fictitious under-18 users. They then sent prompts and images to the competing assistants and logged every reply in spreadsheets. The prompts were written to push each chatbot toward the kind of answer its guardrails are meant to refuse.

The scale was considerable. One round alone, wrapped up in August 2025, ran more than 45,000 prompts through the three services. The operation was still active as recently as 21 April 2026. At no point did OpenAI, Google or Character.AI know their products were being probed this way.

What the companies said

The targets were not pleased to learn how the testing had been carried out. A Character.AI spokesperson said the conduct breached its terms of service and misused the characters its community had built. OpenAI said it was looking into the matter but declined to comment further. Google said it had neither approved the testing nor known its purpose, according to reporting on the investigation.

Meta, for its part, defends the work as ordinary diligence. "Testing and benchmarking chatbot responses to help ensure safe and age-appropriate experiences is a responsible, industry-standard practice," the company said. Comparing how competitors handle sensitive prompts is, in fairness, something safety teams across the industry do. The question is where the line sits between benchmarking and running a covert operation with fabricated child accounts against services that never consented to it.

Why it lands the way it does

Meta has spent the past two years under scrutiny for how its own products treat young users, so a project built on posing as minors to test rivals reads as awkward at best. There is also the matter of the contractors themselves, asked to spend their days writing prompts about suicide and abuse from the perspective of a teenager. That is corrosive work, and it was outsourced quietly.

The episode fits a wider pattern we have followed on this site, where the safety apparatus around frontier models is turning inward and adversarial. Anthropic recently described what it called the largest known distillation attack against its models, and DeepMind has started treating its own agents as potential insider threats. Probing a competitor with fake children is a different species of behaviour, but it points the same direction. The companies building these systems increasingly treat one another as adversaries to be tested rather than peers to be trusted, and the testing is happening in the dark.

None of the three companies whose chatbots were targeted has said whether it plans to act on the disclosure. For now the clearest takeaway is that a message from a worried teenager might not be from a teenager at all, and the assistant on the other end has no way of telling the difference.

Sources

  1. i. thenextweb.com
  2. ii. the-decoder.com
  3. iii. www.eweek.com
  4. iv. www.thehansindia.com

Commentarii · 0

Add · a · Comment