Google used its I/O 2026 keynote on May 20 to push its Gemini family deeper into multimodal territory, introducing a video-generation model called Gemini Omni and making the latest Gemini 3.5 Flash generally available. The pair signals where the company wants the next year of competition to play out, with longer-running agents on one side and grounded video synthesis on the other.
Gemini Omni accepts a mix of images, audio, video and text, then produces video clips the company says are "grounded in Gemini's real-world knowledge." Google highlights physics fidelity as the differentiator, claiming better handling of gravity, fluid dynamics and kinetic energy than earlier video models. Every generated clip carries an embedded SynthID watermark that the Gemini app can verify, a provenance feature Google has been pushing across its image and audio tools since 2024.
What Omni actually does
The pitch sits closer to a world model than a text-to-video toy. Internally the company has spent the past year integrating Project Genie with two decades of Street View imagery, allowing Gemini to build interactive scenes anchored in real places. Omni inherits the same spatial reasoning, which is why Google leans hard on the physics claims rather than aesthetics. Whether it produces footage that holds up under close inspection is something independent reviewers will start probing this week.
Rollout starts Tuesday for Google AI Plus, Pro and Ultra subscribers through the Gemini app and Google Flow, according to CyberNews. YouTube Shorts gets Omni next week. Developers and enterprise API customers wait a few weeks longer, which fits the now-familiar pattern of consumer-first launches with paid tiers gating the headline feature.
Gemini 3.5 Flash, priced for agents
The quieter announcement may matter more for working engineers. Gemini 3.5 Flash is now generally available at $1.50 per million input tokens and $9 per million output, with a one-million-token context window. Google says it runs roughly four times faster than comparable models in the same price band and is tuned for what the company calls long-horizon agent work: code generation that spans multiple files, browser automation, and the slow-running pipelines that have been pushing Gemini in Chrome lately.
The pricing is the interesting part. At those numbers, Flash sits below most reasoning-tier models and well below frontier pricing, which makes it a plausible default for the bulk of an agentic workflow before kicking up to a larger model on harder steps. The model is the first in what Google calls the 3.5 family, suggesting that a Pro-tier release will follow.
The rest of the keynote
Google also showed Pomelli, an agent that helps with brand-book and website setup, and announced Co-Scientist, a Gemini-powered partner pitched at researchers. A $100 AI Ultra plan rounds out the subscription lineup, sitting above the existing AI Pro tier. Antigravity, an internal coding environment trailed in pre-event leaks, got a brief on-stage demo focused on agent orchestration.
For readers tracking the wider field, this is the second major frontier-model push from Google inside a month, following its May rollout of Gemini in Chrome. The cadence reflects a company pricing aggressively to keep developer mindshare while OpenAI's recent GPT-5.2 release and Anthropic's enterprise focus carve up the rest of the market.
Worth watching over the next fortnight: third-party benchmark results for 3.5 Flash, and whether Omni's physics claims survive contact with adversarial prompts from the usual stress-testers. Google has been burned before on demo footage that looked sharper in the keynote than in real hands.
Sources
- i. blog.google
- ii. tech.yahoo.com
- iii. cybernews.com
- iv. thenextweb.com
Commentarii · 0