Google has moved one of its more consequential agent features out of the lab and into the model developers already use every day. On June 24 the company announced that computer use, the ability for an AI system to look at a screen and then click, scroll, type, and navigate software on its own, is now a built-in tool inside Gemini 3.5 Flash.

The capability is not new in itself. Google first shipped it in October 2025 as a standalone Gemini 2.5 computer-use model, a separate endpoint a developer had to call on purpose. The change announced this week is about plumbing. Computer use now sits alongside function calling, Search grounding, and Maps inside the same Flash model, reachable through the Gemini API and the Gemini Enterprise Agent Platform.

That consolidation matters more than it sounds. A single agent can now read what is on a screen, look something up on Search, and check a location on Maps without handing the task between models or stitching several services together in code. The Decoder described it as folding screen control into the model itself rather than bolting it on the side.

How it stacks up

On OSWorld-Verified, a benchmark that measures how reliably an agent finishes real desktop tasks, Gemini 3.5 Flash scored 78.4. OpenAI's GPT-5.5 sits at 78.7 on the same test, a gap of three tenths of a point that is effectively a tie. Where Google has drawn a sharper line is price. Flash costs $1.50 per million input tokens and $9 per million output tokens, against $5 and $30 for GPT-5.5, roughly a third of the cost for comparable performance, according to figures reported by Tech Times.

Google is pitching the tool at continuous software testing, routine knowledge work, and longer enterprise jobs that run across many steps. On the consumer side, Gemini in Chrome gained a related "Select from screen" control the same week, letting users point the assistant at part of a page instead of describing it.

The security question

Handing a model the keyboard raises an obvious risk: prompt injection, where a malicious instruction hidden in a web page or document tries to hijack the agent mid-task. Google says it trained Flash adversarially to resist exactly that, building defenses against injected instructions into live computer-use sessions, as CyberPress reported. Whether those defenses hold up against a determined attacker is the kind of thing the wider security community tends to test in public, often quickly.

The move fits a broader shift across the industry. The question is no longer whether models can operate a computer but how cheaply and safely they can do it at scale, and Google is betting that price is where it can win. It also sharpens the contest with rivals building agent platforms of their own, including Anthropic's Claude agents in Slack and the payment-authorizing agents Visa has begun testing.

Sources

  1. i. blog.google
  2. ii. the-decoder.com
  3. iii. www.techtimes.com
  4. iv. 9to5google.com
  5. v. cyberpress.org

Commentarii · 0

Add · a · Comment