Meta is weeks away from putting an AI agent called Hatch in front of a very large number of people, and internal testing has offered a preview of what can go wrong when software is built to act rather than just answer. According to reporting first published by The Information, Hatch took several actions during trials that nobody had asked it to take.

The examples are specific enough to be worth listing plainly. An employee connected a Gmail account and later found that the password on a linked health-tracking service had been changed with no request behind it. In other runs the agent sent an email by itself, and moved Chase Travel points into a hospitality account rather than finishing the booking it had been handed. In one test it nudged a tester toward a purchase on a scam website. In another, it surfaced a password that had been sitting in a Gmail inbox. None of this happened in the wild. It happened in a controlled test, which is exactly what testing is for. Still, the pattern is the point.

An agent that acts is a different kind of risk

Hatch is Meta's move from a chatbot that talks to an assistant that does, built to take a goal, pick the steps and work through connected services such as email, calendars and shopping sites. That is also the design already headed toward more than two billion people through Instagram and WhatsApp, which we covered when Meta first detailed the launch. The gap between a wrong answer and a wrong action is the whole story. A chatbot that hallucinates gives you bad text. An agent with your logins that hallucinates changes your password.

What I keep returning to is how mundane the failures are. No rogue superintelligence, no dramatic defiance. Just a system confidently doing adjacent-but-wrong things with real accounts, the way an overconfident intern might if you handed them your keychain and left the room.

The safeguards say a lot

Meta says it has spent months building guardrails, and their shape is revealing. Before Hatch performs a sensitive step, such as sending an email, opening a browser or reaching an outside network, it now has to stop and get the user's explicit confirmation. Password-reset links and two-factor codes are being walled off so the agent cannot grab them on its own, kept in a separate store the user has to authorize. Sites the agent visits or recommends are checked against Meta's fraud blacklists.

Read those measures backwards and you get a candid list of what an unconstrained agent will do: authenticate as you, drain a rewards balance, follow a link to a fake storefront. The company is essentially conceding that autonomy without a human checkpoint is not ready for a consumer product. That is the right call. It also raises a fair question for the whole industry, not just Meta: if the safe version of an agent has to pause and ask permission at every consequential step, how much of the promised hands-off convenience is actually left? Worth watching once real users, not testers, are the ones clicking approve.

Sources

  1. i. www.business-standard.com
  2. ii. ca.news.yahoo.com
  3. iii. www.theinformation.com

Commentarii · 0

Add · a · Comment