
Google Unveils Gemini 3.5 Transcribe: Next-Gen Speech Model Filters Filler Words and Cuts Latency by 70%
A New Standard for Voice: Moving Beyond Literal Audio Transcription
Google has officially introduced Gemini 3.5 Transcribe, its most advanced speech-to-text model engineered to convert raw, natural spoken audio into clean, formatted text in real time. Joining the Gemini Audio family alongside Gemini 3.5 Live and Live Experimental, the new release slashes latency by 70% compared to previous generations and natively supports more than 85 languages across developer platforms, Android devices, and desktop environments.
According to the engineering leadership behind the project-including Diego Melendo Casado, Senior Director of Engineering for Gemini Audio, and Chief of Staff Luke Leonhard-the model was built to solve the long-standing flaws of legacy automated speech recognition (ASR) engines. Traditional models often stumble in noisy environments, misinterpret complex technical jargon, and transcribe speech disfluencies literally, creating cluttered and unreadable text that requires tedious manual editing.
Gemini 3.5 Transcribe bridges the gap between acoustic capture and contextual language understanding. Rather than simply logging every syllable, the model processes spoken intent directly, delivering polished transcripts formatted with punctuation, correct capitalization, and inline semantic structure right out of the box.
Smart Transcription: Handling Real-Time Self-Corrections and Hesitations
Spontaneous human speech is naturally disorganized. People hesitate, stutter, repeat themselves, or change their minds mid-sentence. Where conventional transcription software fails by recording every “um” and “ah,” Gemini 3.5 Transcribe applies contextual filtering to refine spoken thoughts into coherent sentences.
The model natively parses mid-sentence self-corrections. For example, if a speaker says, “Let’s schedule the meeting for Tuesday-no, wait, Wednesday morning,” the engine detects the correction and directly outputs: “Let’s schedule the meeting for Wednesday morning.”
Beyond removing filler words, Gemini 3.5 Transcribe supports continuous voice-driven text manipulation. Users can dictate content, issue verbal formatting commands, correct specific terms, or alter writing tone entirely through natural voice interaction without touching a keyboard.
Benchmark Performance: Record-Low Error Rates and 70% Latency Gains
Independent evaluations conducted by Artificial Analysis highlight substantial performance leaps across both real-time streaming and asynchronous batch workflows:
- Word Error Rate (WER): Achieves an average WER of 4.0% in live streaming mode and 2.6% for pre-recorded audio processing.
- Latency Reduction: Time-to-final-transcription improves by 70% compared to Google’s Chirp 3 model released in 2025.
- Multilingual Accuracy: On the FLEURS benchmark across top global languages and locales, the model achieved a 5.50% WER in streaming mode and 5.04% WER in non-streaming tests.
- Alphanumeric Entity Capture: Demonstrates high fidelity in identifying complex alphanumeric strings-such as postal codes, tracking IDs, and serial numbers-even against loud ambient background noise.
Global Language Support, Custom Vocabularies, and Speaker Diarization
To serve international workflows, Gemini 3.5 Transcribe automatically identifies and transcribes more than 85 languages, adapting seamlessly to regional dialects and diverse accents. The engine can also handle dynamic live language switching within a single audio stream without loss of context.
For specialized industries, Google has integrated a Custom Vocabulary engine. Enterprise teams in healthcare, legal, finance, and software engineering can feed proprietary terms, unique product names, and technical terminology directly into the model to guarantee pinpoint spelling and semantic precision.
Additionally, the model features built-in multi-speaker diarization for recorded audio. It can accurately differentiate and attribute dialogue for up to three distinct speakers with word-level timestamps, making it an ideal engine for transcribing podcasts, board meetings, and journalistic interviews (with experimental support for larger speaker groups underway).
Developer Architecture: Dual APIs and Multimodal Function Calling
Google is deploying Gemini 3.5 Transcribe across two dedicated developer interfaces designed for distinct architectural requirements:
- Live API (
gemini-3.5-transcribe-live): Provides continuous, bidirectional streaming with sub-second latency for interactive voice bots, real-time closed captioning, and conversational AI agents. - Interactions API (
gemini-3.5-transcribe): Optimized for batch processing of recorded audio, call center analytics, and post-meeting documentation with precise speaker attribution.
A standout technical capability is Function Calling. Gemini 3.5 Transcribe does not just output text; it can route complex instructions to other Gemini models in the background. In practice, a user can verbally command the system to analyze an open file or generate an illustration, and the transcription layer will parse the intent, invoke the appropriate downstream model, and execute the task autonomously.
Ecosystem Rollout: Android, macOS, Antigravity, and Chrome
Google is rolling out Gemini 3.5 Transcribe across its core consumer and enterprise ecosystem:
- Android & Gboard: Powers the new Rambler feature on devices like the Pixel 11 series, transforming rambling voice notes into structured text while allowing spoken inline editing.
- Gemini App on macOS: Powers the Speak to Window capability, pairing screen context with voice commands to summarize local documents or generate assets at the cursor.
- Google Antigravity & AI Studio: Enables voice-driven app development (“vibe coding”) and contextual code prompting by reading screen state, active documents, and chat history.
- Google Chrome: An upcoming update will bring universal talk-to-type capability across any web field, letting users dictate emails, draft social posts, or query Gemini directly from the browser.
Google confirmed that subsequent rollouts will integrate the transcription model into Google Docs, Gmail, Google Keep, Search Live, and Gemini Live.
Industry Adoption and Enterprise Availability
Major developer platforms and infrastructure providers-including LangChain, LiveKit, Vercel, Agora, Pipecat, Fishjam, and Vision Agents-have already integrated Gemini 3.5 Transcribe to power real-time voice applications.
Early enterprise adopters, such as smartphone maker Vivo, healthcare platform Intellitek Health, and translation service Lingopal, have reported substantial improvements in conversational latency and terminology recognition.
Gemini 3.5 Transcribe is currently available in public preview for developers via Google AI Studio and Google Antigravity, and for enterprise customers through the Gemini Enterprise Agent Platform, with wider availability for customer experience suites slated for later this year.
The Strategic Horizon: Voice as the Primary Computing Interface
The release of Gemini 3.5 Transcribe highlights Google’s broader strategy: transitioning computing away from physical keyboards toward voice as a frictionless, primary input layer. By combining contextual screen awareness, autonomous function calling, and real-time disfluency cleanup, Google is positioning voice not just as an accessibility tool, but as the fastest way to operate complex software in everyday workflows.




