Google DeepMind launches Gemini 3.5 Transcribe, a new speech-to-text model with real-time and pre-recorded transcription APIs
Google released Gemini 3.5 Transcribe, a new speech-to-text model available via two APIs: gemini-3.5-transcribe-live (real-time streaming via Live API) and gemini-3.5-transcribe (pre-recorded audio with speaker attribution and word-level timestamps via Interactions API). It is now in public preview in the Gemini API/Google AI Studio and Gemini Enterprise Agent Platform, integrated into consumer surfaces like Rambler on Android, Gemini app on macOS, Google Antigravity, and coming soon to Chrome and Gemini Enterprise for Customer Experience. It replaces the previous model, Chirp 3.
Entities: Google DeepMind, Gemini 3.5 Transcribe, Gemini API, Google AI Studio, Gemini Enterprise Agent Platform, Rambler
1 primary
What happened
Google DeepMind released Gemini 3.5 Transcribe, a speech-to-text model available through two new API endpoints: one for real-time streaming and one for pre-recorded audio with speaker attribution and word-level timestamps. It is in public preview across the Gemini API, Google AI Studio and Gemini Enterprise Agent Platform, and is already integrated into consumer products including Rambler on Android, the Gemini macOS app and Google Antigravity. It replaces Chirp 3 as Google's transcription model.
Why it matters
Developers building voice or transcription features now have a new default Google option to evaluate against Whisper, Deepgram, AssemblyAI and ElevenLabs, with third-party-measured error rates (4.0% streaming, 2.6% non-streaming) that are genuinely competitive if they hold up in practice. Enterprises using Google's stack get an immediate upgrade path since the model is already live in several products, but anyone planning production deployment should note this is preview software with no published pricing or rate limits yet. The broader significance is limited: this is an incremental capability upgrade within Google's existing ecosystem, not a shift in who controls voice AI infrastructure.
What is noise
The "most precise speech-to-text model yet" framing and the 70% latency improvement are Google's own claims, sourced partly from Artificial Analysis but not independently replicated here, and preview-stage performance often changes before general availability. The long list of partner and integration names (Agora, LangChain, LiveKit, Pipecat, Vercel, vivo, etc.) reads as ecosystem packaging to signal momentum rather than evidence of proven real-world adoption at scale.
Watch next
- 01Independent WER and latency benchmarks (e.g. from Artificial Analysis or third-party developers) replicating the claimed 4.0%/2.6% error rates and 70% latency improvement outside Google's own post
- 02Pricing and rate limits once the API exits public preview, since cost will determine whether it actually displaces Whisper, Deepgram, AssemblyAI or ElevenLabs in production pipelines
- 03Whether multi-speaker attribution beyond three speakers and the promised Chrome and Gemini Enterprise for Customer Experience integrations ship, or remain stuck as experimental/coming soon
Coverage
1 storyMore capability signals
Full feed →- AI systems outperform expert humans in persuasive communication22 Jun 202681
- Benchmark results show significant improvement in AI agent performance on WorkBench15 Jun 202679
- Introduction of Stateful ReAct Agents for Token-Efficient Autonomous Experimentation16 Jun 202678
- Study reveals flaws in LLM-as-judge safety evaluations due to temperature settings26 Jun 202677