Launch Video Library · AI · Feature · 2026

Gemini 3.8 Text-to-Speech

Google DeepMind introduced custom voice generation and verified cloning capabilities to Gemini 3.8 text-to-speech.

What Google DeepMind shipped

Google DeepMind expanded the capabilities of Gemini 3.8 text-to-speech, introducing tools to design custom vocal personas from scratch and clone existing voices using a 30-second audio sample.

The update targets developers building gaming environments, audiobooks, and podcasts, emphasizing natural language direction for pacing, acting cues, and dialect shifts. It also integrates SynthID watermarking and C2PA credentials to address consent and security in voice synthesis.

How the motion works

When showcasing a purely audio-driven feature like Gemini 3.8 text-to-speech, motion designers often rely on kinetic typography to anchor the viewer. Syncing text animations to the generated voice—using scale or weight shifts to mirror pacing and acting cues—gives invisible audio a physical presence on screen.

To illustrate the transition from natural language prompts to complex vocal outputs, a tight sequence of match cuts can bridge the interface and the result. Cutting from a typed prompt directly to an audio waveform or a stylized representation of the persona creates a clear cause-and-effect relationship without requiring lengthy UI walkthroughs.

For a 60-second spot demonstrating back channeling and dialect shifts, cutting-on-the-beat of the generated speech keeps the momentum high. Quick hard cuts between different vocal personas—gaming characters, podcast hosts, audiobook narrators—establish the model's range efficiently.

When communicating trust and safety features, pacing must shift. Allowing a deliberate hold-time on the final SynthID and C2PA credential screens ensures the consent verification message lands clearly before the concluding call to action.

What to steal from it

  • Anchor audio features with kinetic typography to give invisible products a visual footprint.
  • Use match cuts to compress the journey from text prompt to generated output.
  • Cut on the beat of the voiceover to establish a natural rhythm for the edit.
  • Reserve hold time for critical security or compliance features to ensure they register.

Want a video like this for your product?

Impractical cuts one from a single prompt — the same motion craft, in about twenty minutes.