Google has put two new text-to-speech models into its developer tools, and the headline feature is how little audio they need to imitate a real person. Released on 23 September in the Gemini API and AI Studio, Gemini 3.8 Flash TTS and its lighter sibling, Flash-Lite TTS, can rebuild a consistent vocal profile from a 30-second sample. Google frames this as replicating your own voice, "or a voice the user has rights to use," a caveat that is doing a lot of quiet work.
Beyond cloning, the models are built for control. A developer can describe a voice in plain language, setting its role, accent, and character, and get something usable back across more than 100 languages and dialects. Google has grown its original set of 30 voices into a library of over 2,000, with regional flavours such as Mexican Spanish, Quebec French, and Scots English. Delivery can be steered line by line, and the models can voice back-and-forth conversations rather than a single flat reader.
The two tiers split along cost and coverage. Flash TTS handles 130 languages and ships with 30 prebuilt studio voices, including ones Google names Kore and Puck. Flash-Lite covers 101 languages with the same voice-design surface and is pitched as the high-volume, cheaper option. A remixing feature that tweaks timbre, pitch, pace, and accent on existing voices is listed as coming soon.
The part worth watching
Thirty seconds is not much. It is a voicemail greeting, a snippet of a podcast, a clip lifted from a video call. Google's answer to the obvious misuse worry is contractual rather than technical: the terms say you must have rights to the voice you copy. That places the burden on whoever presses the button, and it is the same line most voice vendors now draw. Whether it holds depends entirely on enforcement, which is hard to see from the outside.
Set against the rest of the field, this is less a leap than a steady tightening of the screws. Rivals have been racing on expressive, low-latency speech all year, including OpenAI, whose voice model recently learned to listen and talk at the same time. Google's contribution is breadth and ease: more languages, more ready voices, and a cloning step short enough that the friction is nearly gone. For developers building narration, dubbing, or assistants, that is useful. For anyone thinking about consent and audio deepfakes, the shrinking sample size is the number to keep an eye on.
Sources
- i. blog.google
- ii. www.unite.ai
- iii. www.marktechpost.com
Commentarii · 0