Google Releases Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking Real-Time Audio Models
On September 15, 2026, Google released two native end-to-end real-time audio models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. Gemini 3.8 Live supports 97 languages, real-time visual context, and asynchronous tool and API calls. The Extended Thinking version adds background multi-step reasoning and ranked first on Artificial Analysis's Speech to Speech Quality Index with a score of 82.6. Both models are available through the API and AI Studio.
On September 15, 2026, Google released two real-time audio models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. Both are native end-to-end speech models that can call tools, process visual information, and carry out complex tasks without interrupting a conversation. According to information published by Google, Gemini 3.8 Live supports real-time visual context and multilingual conversation, while Gemini 3.8 Live Extended Thinking adds deep reasoning capabilities, completing multi-step tasks in the background while the user continues to speak.
Gemini 3.8 Live supports 97 languages and can automatically identify and switch languages during a conversation while maintaining consistent voice performance. Logan Kilpatrick, software lead for Google AI Studio, said the new model supports 97 languages and adds capabilities including asynchronous tool calling. The model can execute tool and API calls asynchronously in the background while maintaining a continuous conversation; users do not need to wait for a task to finish before continuing to speak, and the model can keep talking while the task runs in the background, incorporating the result into the conversation once it is complete. The model also supports real-time visual context, allowing users to ask questions by voice and have the model understand what it is seeing. Scenarios demonstrated by Google include real-time employee training, troubleshooting, and having AI participate in a chess game through visual information. Google said the new model has improved handling of alphanumeric information, including confirmation codes, claim numbers, and technical data.
When developers call the model through the Live API, audio input is priced at USD 0.005 per minute and audio output at USD 0.018 per minute.
Gemini 3.8 Live Extended Thinking can perform deeper reasoning while maintaining a real-time conversation, with background extended reasoning that allows it to think while speaking. The model can inform users that it has begun processing a task through voice prompts such as let me look that up, and execute multi-step tasks in the background while maintaining foreground conversation. Cases demonstrated by Google include completing multi-step bookings through asynchronous function calls, generating React components in real time based on user voice and sketches, and producing business plans and marketing toolkits through natural language.
Data published by Google show that Gemini 3.8 Live Extended Thinking ranked first on Artificial Analysis's Speech to Speech Quality Index with a score of 82.6. It achieved a completion rate of 68.6% on the τ-Voice agent task test and 35.1% on Sierra's τ-Voice-banking test, and scored 97.7% on the Big Bench Audio test. The standard Gemini 3.8 Live ranked second in the Speech Agent Arena.
Gemini 3.8 Live Extended Thinking has also been introduced into Google Workspace, where it can be used in products including Docs, Gmail, and Keep. Google said all audio generated by its AI products will carry an invisible SynthID watermark embedded directly in the audio output to identify AI-generated content.
Why this event matters
The event has a measured impact on 6 industrys. The strongest current signal is positive for Artificial Intelligence, with intensity 85/100 and 85% confidence over a short term horizon.
Artificial Intelligence
- Direction
- positive
- Intensity
- 85
- Confidence
- 85%
- Horizon
- Short term
Cloud Services & Data Centres
- Direction
- positive
- Intensity
- 65
- Confidence
- 70%
- Horizon
- Short term
General Software & IT Services
- Direction
- mixed
- Intensity
- 55
- Confidence
- 60%
- Horizon
- Medium term
Enterprise Software
- Direction
- positive
- Intensity
- 55
- Confidence
- 60%
- Horizon
- Medium term
Robotics
- Direction
- positive
- Intensity
- 50
- Confidence
- 50%
- Horizon
- Medium term
Semiconductor Value Chain
- Direction
- positive
- Intensity
- 45
- Confidence
- 50%
- Horizon
- Long term
Impact figures are analytical estimates that combine direction, intensity, confidence and event importance. They are not investment advice.