Google unveiled Gemini 3.5 Transcribe, its most precise speech-to-text model to date, designed to handle background noise, clean up disfluencies, and recognize specialized jargon across more than 85 languages. The AI audio model debuts on macOS and Android devices while rolling out developer APIs via Google AI Studio.
Google Unveils Gemini 3.5 Transcribe as Its Most Precise Speech-to-Text Model
Heather Diehl reported that Google revealed Gemini 3.5 Transcribe, which the tech giant said is its most precise speech-to-text model to date. Google officially introduced the AI-powered speech-to-text model built to transform unstructured, messy speech into polished and formatted text across multiple platforms. Joining the broader Gemini Audio family, the new release marks a significant upgrade over the previous transcription engine known as Chirp 3, particularly in multilingual performance, speed, and overall accuracy.
Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text,
according to the company. The system is engineered to capture natural speaking styles while interpreting user intent and adapting to specialized terminology in real time. Unlike conventional speech recognition models that often stumble through complex industry jargon, heavy background noise, and natural verbal hiccups, Gemini 3.5 Transcribe converts raw audio directly into ready-to-use output for emails, messages, and documents.
Performance Metrics and Smart Disfluency Cleanup
Under the hood, the model delivers measurable gains in speed and error reduction. Google states that the new AI model is roughly 70 percent faster from raw voice input to final transcribed text compared to Chirp 3. Independent benchmarks measured by Artificial Analysis show that the system achieves an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming use-cases.

For live-speech tasks, Google measures the error rate at 5.5 percent, representing an improvement over Chirp 3, which Google measures at 7.32 percent. Beyond raw transcription accuracy, the model handles stream-of-consciousness speech by automatically stripping out filler words such as ums and ahs and seamlessly executing self-corrections such as let’s meet Tuesday—no, Wednesday. The model also shows strong performance across noisy, real-world environments, accurately capturing alphanumeric entities like postal codes and order IDs.
Ecosystem Rollout on macOS and Android Devices
The rollout begins today in English for all macOS Gemini app users and the Rambler dictation feature on Android in select countries and languages. On macOS, users can long-press the Fn key to start speaking naturally, and the app transcribes their words directly at the cursor position in any application. The Gemini app requires macOS Sequoia 15.0 or later, at least 8 GB of RAM, and approximately 200 MB of installation space.

On mobile devices, the model powers the Gboard Rambler feature on the Pixel 11, across products like the Gemini app and on Android, and in the Gemini app on macOS.
Developer APIs and Advanced Audio Capabilities
Developers can build similar capabilities to plug seamlessly into developer workflows, whether building voice agents, real-time captioning tools, or post-call analytics pipelines.
The model is available across two separate APIs:
- Real-time streaming: Delivers continuous, bidirectional streaming with sub-second latency for interactive voice apps via the Live API using gemini-3.5-transcribe-live.
- Pre-recorded audio processing: Transcribes recorded audio, meetings, call logs, and more with speaker attribution and word-level timestamps via the Interactions API using gemini-3.5-transcribe.
The system also supports multi-speaker identification, capable of attributing speech to up to three distinct speakers in pre-recorded audio with word-level timestamps (support for 3+ speakers is experimental). Furthermore, custom vocabulary features recognize specialized jargon and unique spellings by seamlessly adapting transcriptions to user-provided custom vocabulary. Automatically detecting and transcribing over 85 languages, the model handles regional accents and diverse dialects while supporting function calling to delegate complex tasks such as image generation and file analysis to other Gemini models.
Keep reading