Gemini 3.8 Flash TTS: How to Use It + Prompt Template

Direct answer: Google launched two Gemini 3.8 text-to-speech models on September 23, 2026. Use Gemini 3.8 Flash TTS when voice fidelity, acting nuance, difficult pronunciation or regional dialects matter most. Use Gemini 3.8 Flash-Lite TTS for lower-latency, high-volume work such as read-aloud features and voice-agent pipelines.[3][6][7]

You can test both models in Google AI Studio now. Developers can use them through the Gemini API, while Gemini Enterprise support is listed as coming soon. Flash TTS is also rolling out in Gemini Notebook; Flash-Lite TTS is rolling out in Google Vids.[3]

Quick start: Open Google AI Studio’s speech-generation workspace, select gemini-3.8-flash-tts or gemini-3.8-flash-lite-tts, choose or design a voice, paste the prompt template below, generate a short sample, and only then expand to a long script.

Gemini 3.8 Flash TTS vs Flash-Lite TTS

Choose Best for Languages Model ID
Gemini 3.8 Flash TTS Audiobooks, studio narration, complex two-person dialogue, demanding accents and expressive acting 130 gemini-3.8-flash-tts
Gemini 3.8 Flash-Lite TTS High-volume production, read-aloud tools, voice-agent cascades and everyday single-speaker audio 101 gemini-3.8-flash-lite-tts

Google describes Flash as the creative-quality tier and Flash-Lite as the throughput, latency and cost-efficiency tier. Both accept text and return audio, support an 8,192-token input limit and 16,384-token output limit, and support Batch, Flex and Priority inference.[6][7]

The practical choice is simple:

  • Start with Flash-Lite for routine narration, accessibility audio, product read-aloud and large batches.
  • Move to Flash when a scene needs stronger character identity, subtle emotion, hard pronunciation or a specific regional delivery.
  • Test the same 20–30 second passage on both models before committing a long project.

What is new in Gemini 3.8 TTS?

The launch moves beyond a small set of static voices. Google says Flash TTS can create custom voices with natural-language descriptions, offers more than 2,000 production-ready voices, and supports voice replication from a 30-second sample when the user has the right to use that voice.[3]

Both models support line-by-line performance direction, long-form generation and native two-speaker staging. Google also says generated audio receives SynthID watermarking, while replicated voices require consent verification and include C2PA credentials.[3]

There are important limits:

  • TTS is for controlled speech from a script; Gemini Live is the separate product for interactive, open-ended voice conversations.[8]
  • The model pages list audio generation and caching, but not tools such as Search grounding, function calling, URL context or code execution.[6][7]
  • Voice replication in AI Studio is not available in Illinois, Texas, the EEA, the UK, Switzerland or India at launch.[3]
  • Availability is rolling out, so an account may not show every option immediately.[3]

How to use Gemini 3.8 TTS in Google AI Studio

1. Open the speech-generation workspace

Go to Google AI Studio speech generation. Testing in AI Studio is the fastest route because you can compare voices and revise direction before building an API workflow.[8]

2. Select the correct model

Choose:

  • gemini-3.8-flash-tts for maximum performance quality; or
  • gemini-3.8-flash-lite-tts for a faster, more economical production path.

Do not select the ordinary gemini-3.8-flash text model by mistake. The TTS endpoints have -tts in their model IDs.[6][7]

3. Start with a short audition script

Use 50–100 words that contain the hardest parts of your real project: names, acronyms, numbers, emotional changes and any regional vocabulary. A short audition makes problems cheaper and easier to fix than regenerating a full chapter.

4. Add direction, not just adjectives

Google’s prompting guide recommends defining an audio profile, scene and director’s notes. Useful direction includes the character’s role, environment, emotional state, accent and pacing.[8]

Instead of:

Read this in an exciting voice.

Use:

AUDIO PROFILE
A knowledgeable technology host speaking to curious beginners.

SCENE
A quiet, treated podcast studio. The listener is following a practical tutorial.

DIRECTOR'S NOTES
Style: Warm, confident and helpful; never salesy.
Pacing: Medium pace. Slow slightly before commands and model names.
Accent: Neutral international English.
Pronunciation: Say “TTS” as individual letters.

TRANSCRIPT
Google has released two Gemini 3.8 text-to-speech models. Here is how to choose the right one.

5. Compare, revise and save the winning setup

Generate several short versions. Compare pronunciation, consistency, breath noise, pauses and emotional fit. Change one variable at a time—voice, style, pacing or script—so you know which adjustment improved the result.

Copy-paste prompt template

AUDIO PROFILE
Name/role: [Who is speaking?]
Audience: [Who is listening?]
Core voice traits: [warm, authoritative, playful, calm, etc.]

SCENE
Location: [studio, classroom, help desk, story world]
Context: [what is happening and why]
Emotional atmosphere: [reassuring, urgent, reflective, energetic]

DIRECTOR'S NOTES
Style: [specific performance direction]
Pacing: [slow/medium/fast plus where to pause]
Accent or dialect: [be precise, but only request accents you actually need]
Pronunciation: [names, brands, acronyms and phonetic hints]
Energy: [low/medium/high and where it changes]

TRANSCRIPT
[Paste the exact words to speak.]

For a two-person scene, label each line consistently:

Make HOST calm and clear. Make GUEST energetic but natural.

HOST: What changed in Gemini 3.8 text to speech?
GUEST: You can now direct the performance more precisely and choose between a creative-quality model and a high-throughput model.

Gemini TTS supports single-speaker and multi-speaker generation; Google’s current guide limits a multi-speaker configuration to two speakers.[8]

Useful audio tags

Google’s TTS guide documents inline tags such as [whispers], [laughs], [sighs], [gasp], [excited], [serious], [tired] and [shouting]. Tags can change delivery for a line or section without rewriting the whole prompt.[8]

Example:

[serious] Back up the original recording before you edit it.
[short pause]
[calm] Then test the new voice on a short passage.

Treat tags as creative controls, not guaranteed commands. Google says there is no exhaustive list of tags that always work, so test each one against your chosen voice and language.[8]

Migration checklist from Gemini 3.1 Flash TTS Preview

Flash-Lite is Google’s recommended replacement for gemini-3.1-flash-tts-preview when the priority is high-throughput production.[7]

Before switching production traffic:

  1. Replace the old model ID with gemini-3.8-flash-lite-tts or gemini-3.8-flash-tts.
  2. Re-test your speaker configuration and style instructions because the 3.8 model pages describe a revised schema.
  3. Check output handling: the migration notes list WAV as the default output and also reference L16, μ-law and A-law options.[6][7]
  4. Re-run pronunciation tests for names, numbers, abbreviations and multilingual passages.
  5. Compare latency and output consistency on a representative batch.
  6. Keep the old production path available until the new output passes listening QA.
  7. Review the live pricing page before scaling. At the time of this check, Google’s public pricing page had not yet added a clearly labelled Gemini 3.8 TTS section, so this article does not quote an unverified price.[9]

Voice-replication safety checklist

A technically possible voice clone is not automatically safe or lawful to publish. Use this checklist:

  • Obtain explicit permission from the person whose voice will be replicated.
  • Keep a record of the consent and the exact approved use.
  • Do not imitate a celebrity, customer, employee or private person without authorization.
  • Label synthetic audio when a listener could reasonably mistake it for a real recording.
  • Store reference recordings with restricted access and delete them when they are no longer needed.
  • Test whether the generated voice could create a misleading endorsement or statement.

Google says its replication flow includes consent verification and that generated Gemini Audio clips are watermarked with SynthID.[3] Those safeguards support responsible use, but they do not replace permission, disclosure or human review.

Practical QA checklist before publishing audio

Listen to the final file from beginning to end and check:

  • names, acronyms, dates, currency and URLs;
  • speaker separation in two-person scenes;
  • unwanted laughs, breaths, pauses or emotional shifts;
  • consistent accent and character identity across sections;
  • clipped openings or endings;
  • factual accuracy of the script itself;
  • disclosure and consent requirements;
  • loudness and format requirements for the destination platform.

For long-form work, generate in sections and keep a locked voice prompt. This makes it easier to replace one bad paragraph without changing an entire chapter.

FAQ

Is Gemini 3.8 Flash TTS available now?

Google says both Gemini 3.8 TTS models began rolling out on September 23, 2026 in Google AI Studio and the Gemini API. Enterprise API availability is listed as coming soon.[3]

Which Gemini TTS model should I use?

Use Flash for maximum fidelity, acting nuance and dialect coverage. Use Flash-Lite for high-volume, low-latency and cost-sensitive workloads.[6][7]

Can Gemini 3.8 TTS clone a voice?

Google says Flash TTS can replicate a voice from a 30-second sample when the user has the rights to that voice. Consent verification is part of the feature, and regional restrictions apply in AI Studio.[3]

How many speakers can Gemini TTS generate?

Google’s speech-generation guide currently supports single-speaker audio and configurations with up to two speakers.[8]

Does Gemini 3.8 TTS support Urdu, Arabic and Hindi?

The Flash model’s official language table includes Hindi, Standard Arabic, Egyptian Arabic and multiple related regional languages. The table lists Sindhi but does not explicitly list Urdu on the 3.8 model page, so verify your exact language in AI Studio before planning production.[6]

What does Gemini 3.8 TTS cost?

No clearly labelled 3.8 TTS price was present on Google’s public Gemini API pricing page when checked on September 23, 2026. Check the live pricing page before sending production traffic; do not assume the ordinary Gemini 3.8 Flash text-token price applies to generated audio.[9]

Sources

Leave a Comment

muddaser logo

Public Speaker, Softskills trainer and technology enthusiast

Contact

Muddaser Altaf

Social Address