Notice Nearby

Company newsroom — sourced from Google · Notice Nearby

Build real-time voice applications with Gemini 3.8 Live and 3.5 Transcribe

Google

September 15, 2026

Read original on the Google newsroom

Build real-time voice applications with Gemini 3.8 Live and 3.5 Transcribe Sep 15, 2026 | x.com Facebook LinkedIn Mail Copy link New Gemini Audio models are available for developers to build more intelligent conversational experiences via the Gemini API and Google AI Studio. Alisa Fortin Product Manager, Google DeepMind Thor Schaeff Member of the Technical Staff (DevX), Google DeepMind x.com Facebook LinkedIn Mail Copy link . Inlining them here makes them available in the DOM for the page. Your browser does not support the audio element. Listen to article [[duration]] minutes This content is generated by Google AI. Generative AI is experimental Voice Speed Voice Speed 0.75X 1X 1.5X 2X Today, we released new Gemini Live models in the Gemini API and Google AI Studio, expanding our developer suite for building real-time, voice-first product experiences: Gemini 3.8 Live and 3.8 Live Extended Thinking: Gemini 3.8 Live brings a step change to our native speech-to-speech models, capable of performing tasks while maintaining dialogue. For complex requests, 3.8 Live Extended Thinking delivers deeper reasoning, ranking #1 on Artificial Analysis’ Speech-to-Speech leaderboard. Gemini 3.5 Transcribe: Our dedicated speech-to-text model brings highly precise transcription across 85+ languages. Released last month, it achieved an average Word Error Rate (WER) of 4.0% (streaming) and 2.6% (non-streaming). Gemini 3.8 Live & 3.8 Live Extended Thinking: Build more intelligent conversational agents Our new models, Gemini 3.8 Live and 3.8 Live Extended Thinking enable developers to build voice agents that can reason and execute tasks while maintaining the flow of conversations. Key capabilities include: Asynchronous function calling: Execute API and tool calls in the background while continuing to stream audio responses to the user Visual context: Ground dialogue in live visual inputs to help enable agents that can understand what users say and see Alphanumeric precision: Accurately parse confirmation codes, claim numbers, and technical data Multilingual support: Reach global audiences with coverage for 97+ languages and accent consistency Incremental content updates: Seamlessly merge real-time audio with structured data to return context-aware responses 3.8 Live Extended Thinking also supports configurable thinking to help handle complex, multi-step reasoning in the background, while responding or narrating its progress in the main conversation. These models represent a step-change from our previous live models and provide a more streamlined alternative to cascaded architectures. Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are available via the Live API. Competitively priced at $0.005/min for audio input and $0.018/min 1 for audio output, they allow developers to scale voice applications with industry-leading performance. Developers can also access the models through Agora, Fishjam, LiveKit, LangChain, Pipecat, Vercel, and Vision Agents, our Live API integration partners that handle media streaming infrastructure for real-world deployment: Gemini 3.5 Transcribe: Convert streamed speech to text Real-time speech understanding is critical for voice-first interfaces. Last month, we released Gemini 3.5 Transcribe for low-latency transcription with high precision, achieving a 4.0% WER, and useful features: Automatic code-switching: Handle intra-sentence and inter-sentential code- and language-switching without manual configuration Custom vocabulary biasing: Steer speech recognition toward domain-specific terms, uncommon jargon, company names, and proper nouns by passing a custom_vocabulary list of up to 1,000 terms Smart transcription mode: Deliver polished, reader-ready transcripts with structured formatting, self-corrections, and disfluency removal that eliminates filler words 3.5 Transcribe supports 85+ languages and provides a strong listening engine for voice experiences and stateless tasks like sub-second captioning, call center agents, and real-time audio analytics. You can also access the model via the Interactions API to transcribe audio files up to 1 hour long with structured timestamps and speaker labeling. Read our developer guide to learn more. Our complete audio suite for developers To get started, try out the models in ai.studio/live, clone example apps from GitHub, or equip your agent with our live api skill. You can also create audio experiences with our speech and music generation models, all available in the Gemini API: Gemini 3.5 Live Translate: Speech-to-speech translation across more than 70 languages Gemini 3.1 Flash TTS: Highly configurable speech generation (with more updates coming soon) Lyria 3.5: Production-grade music generation The mic is yours, and we can’t wait to hear what you build! Get the latest news from Google in your inbox Sign up for our newsletters with product updates, event information, special offers, and more. Done. Just one step more. Check your inbox to confirm your subscription. You can also subscribe with a different email address. Your information will be used in accordance with Google's privacy policy. You may opt out at any time. Posted in: 1 *Estimate based on $3/1M tokens for input and $12/1M tokens for output Related stories Developer tools DevFest is back By Justyna Politanska-Pyszko & Natalie McHugh AI The latest AI news we announced in August 2026 By News from Google Team Developer tools Pairing Google Antigravity with Gemini 3.7 Flash solves notable multi-agent math and engineering problems. Developer tools Gemini Omni 1.1 Flash lets you build with more control By Anish Nangia & Alisa Fortin Developer tools How developers build AI for good with Gemma 4 By Glenn Cameron & Kristen Quan Developer tools Inside the Gemmaverse: Celebrating one billion Gemma downloads By Clement Farabet & Olivier Lacombe — Company newsroom — sourced from Google. Matter furnished by the company. Not a Notice Nearby paid placement. This page reprints matter furnished by the company from its official newsroom. Notice Nearby did not write this release. Read the original: https://blog.google/innovation-and-ai/technology/developers-tools/build-real-time-voice-applications-gemini-audio

Company newsroom — sourced from Google. Matter furnished by the company. Not a Notice Nearby paid placement. NN-PR-NR-2026-0263. This page reprints matter furnished by Google from its official newsroom. Notice Nearby did not write it, and it is not a $79 paid placement. It is not a legal public notice, not an obituary, and not an official Notice Nearby announcement. Record of Sale, LLC · Oregon.

The chain stores a hash, not the notice. The hash is not statutory publication. View on blockchain. Base stores the content hash, publication number, press-release id, and timestamp — not the full release text. The complete release remains on Notice Nearby. This is advertising, not a legal notice.

9111d335 5acab58d d098a53b 3fc693fb 7105e01d 86dd0396 21beeea0 1633d275

All press releases · Official newsroom