Google Launches Gemini 3.8 Flash TTS With Voice Design and Voice Replication
On September 22, 2026, Google made two new text-to-speech models generally available in the Gemini API: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. The Gemini API release notes for that date list the models alongside a new Voices endpoint, voice design, voice replication and an extended voice library.
Text-to-speech turns written text into spoken audio. Businesses use it for voice assistants, in-app narration, phone systems, training content and accessibility features. If you are planning voice features, whether through Android app development, an iPhone app or a web application, this release changes what is possible and, just as importantly, where some of it is allowed.
The headline feature, voice replication, carries consent requirements and regional limits that affect UK companies in particular. This article covers what is available now, what is still to come and what to check before building.
What Google released on September 22, 2026
Google described the two models in a blog.google post published on September 23, 2026. Gemini 3.8 Flash TTS is the higher-quality model, aimed at expressive, directed speech. Gemini 3.8 Flash-Lite TTS is positioned for high-volume, cost-sensitive use.
According to that post, both models are available now in the Gemini API and Google AI Studio. Gemini 3.8 Flash TTS is also in Gemini Notebook and Gemini 3.8 Flash-Lite TTS is in Google Vids. Access through Gemini Enterprise is described as coming soon, so companies that buy Google AI through that route do not have it yet. A voice remixing feature, for adjusting the timbre, pitch, pace and accent of a library voice, is also described as coming soon.
Google's speech generation documentation gives the model IDs as gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts and states that both support single-speaker and multi-speaker output, voice design and voice replication.
Three ways to get a voice
Choosing from the voice library
The documentation describes a new endpoint for listing voices, which can be filtered by language, region, accent, gender, pitch, persona and context. Google's figures for the size of the library differ between its own pages: the release notes mention more than 150 voices, while the blog post refers to more than 2,000. Both describe an expansion from the 30 original prebuilt voices. Check the endpoint itself for the voices available to your project and language.
Voice design
Voice design creates a new synthetic voice from a written description. The blog post describes customizing role, accent and voice characteristics in natural language. For a brand that wants a consistent voice across an app, a phone line and video content without hiring and re-booking a voice actor, this is the lowest-risk option, because the voice does not belong to a real person.
Voice replication
Voice replication creates a synthetic copy of a real person's voice. Google's voice replication documentation says the process needs two recordings from the same adult speaker: a reference clip of 10 to 30 seconds of clean speech, and a consent clip in which the speaker reads a set statement confirming that they own the voice and consent to Google creating a synthetic voice model from it. The documentation provides that statement in 30 languages.
Consent checks, watermarking and regional limits
Google's blog post states that a voice cannot be created unless the verbal consent recording matches the reference speaker. It also states that every audio clip generated by its Gemini audio models is watermarked with SynthID, Google's system for marking AI-generated content, and that voice replication is backed by C2PA content credentials.
The regional limit is the detail most likely to affect readers of this blog. The post states that voice replication through AI Studio is not available in Illinois, Texas, the European Economic Area, the United Kingdom, Switzerland and India. The statement refers to AI Studio specifically. We did not find a Google page that sets out availability for replication through the API in those regions, so a UK business, or a US business with users in Illinois or Texas, should confirm the position with Google before planning a feature around it.
Separately from Google's own rules, copying a person's voice can raise questions under privacy, biometric data, publicity and employment law, and those rules differ between the US and the UK and between US states. This article is not legal advice. Take advice before recording staff, customers or voice actors for replication.
Technical limits worth knowing
The speech generation documentation lists several limits that shape a design. The models accept text only and return audio only, so they do not listen or hold a conversation; real-time, two-way voice is handled by Google's separate Live models. Multi-speaker generation is limited to two speakers in a single request. Streaming requests return raw audio chunks, which lets an app start playback before the whole clip is ready.
On languages, the documentation states that Gemini 3.8 Flash TTS supports over 130 languages and Gemini 3.8 Flash-Lite TTS over 100, with a table showing support for each.
Custom voices have storage rules. According to the voice replication documentation, a stored voice is kept for one year, with a limit of 200 custom voices per project. A voice held as a client-managed key instead expires after seven days. A product that lets each end user create a personal voice would reach a 200-voice limit quickly, so the storage model needs to be decided early.
What this means for your business
What follows is our interpretation, offered as guidance.
For most businesses, the practical gain is in the library and voice design, not replication. A support line, an onboarding walkthrough, a product video or an accessibility read-aloud feature can now use a distinctive, consistent voice in many languages without a recording studio. These uses involve no real person's voice, which keeps the consent and regional questions out of the picture.
Voice replication suits narrower cases, such as a founder or presenter who wants their own voice on content they do not have time to record. It needs a documented consent process of your own in addition to Google's check, and a plan for what happens when that person leaves the company or withdraws consent.
It also helps to be clear about what this release is not. A text-to-speech model gives a product a voice. It does not make it a voice agent. If you are weighing a talking assistant against a simpler automation, our comparison of AI agents, chatbots and workflow automation explains the difference.
What to do next
- List the places in your product where spoken audio would help users, and decide whether each needs pre-generated clips or live streaming.
- Test both models in Google AI Studio with your real scripts, in every language you serve, and have native speakers judge the results.
- Check the Gemini API pricing page for current rates. Audio output is priced separately from text, and the page shows introductory rates with an end date, so estimate costs on the standard rate.
- If you want voice replication, confirm regional availability with Google and take legal advice before collecting any recordings.
- Tell users when they are hearing a synthetic voice. Being open about it is good practice and avoids surprises.
Conclusion
Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS have been generally available in the Gemini API since September 22, 2026, bringing a larger voice library, voice design from a text description and consent-gated voice replication. Gemini Enterprise access and voice remixing are announced but not yet released, and Google says replication through AI Studio is unavailable in the UK, the EEA, Switzerland, India, Illinois and Texas.
Start with the low-risk options, test with your own content and treat replication as a separate decision with its own legal review. Entrant Technologies builds websites, web applications, mobile apps and custom software, and if you want to scope a voice feature for your product, you can request a quote.