Meta releases Muse Voice Transcribe, a real-time speech transcription model with built-in speaker diarization and sentence detection
Meta's Superintelligence Labs released Muse Voice Transcribe, a real-time audio model that transcribes speech, detects sentence boundaries, and distinguishes up to 20+ speakers in a single model without separate systems. It processes audio in 80ms chunks with adaptive per-word delay, supports 70+ languages (25 tested in depth), handles code-switching, and is now available via Meta AI, Muse Code, and the Meta Model API at 0.18 dollars per hour (3 dollars per 1000 audio minutes).
Entities: Meta, Meta Superintelligence Labs, Muse Voice Transcribe, Muse Spark, Meta AI, Muse Code
0 primary
What happened
Meta's Superintelligence Labs shipped Muse Voice Transcribe, a real-time speech-to-text model that also detects sentence boundaries and distinguishes 20+ speakers in one system. It processes audio in 80ms chunks, supports 70+ languages (25 tested in depth), handles code-switching, and is generally available now via Meta AI, Muse Code and the Meta Model API at $0.18 per hour ($3 per 1,000 audio minutes). Independent benchmarks from Artificial Analysis reportedly show 3.1% word error rate at 0.16s latency, ahead of ElevenLabs (3.6%) and AssemblyAI (4.0%).
Why it matters
This is a real, priced, generally-available product that developers and enterprises can procure today for live transcription, captioning, meeting tools or voice agents, so it's immediately actionable for anyone comparing streaming speech-to-text vendors. The combination of low latency, competitive accuracy and low cost pressures existing providers (OpenAI, ElevenLabs, AssemblyAI, Cartesia, Deepgram) on price and forces a re-evaluation of build-vs-buy decisions for speaker-attributed transcription. That said, the impact is bounded to a specific technical niche rather than a shift in AI capability overall.
What is noise
The framing that this is "the foundation for AI assistants that never stop listening" and Zuckerberg's "personal superintelligence" narrative are vendor-supplied speculation layered onto what is otherwise a concrete, incremental product release. Streaming ASR with diarization is already a crowded, competitive commodity market with comparable recent launches from several rivals, so this is best read as a strong price and accuracy move, not a breakthrough. The extraction also lacks primary source links, so the benchmark figures should be verified against Artificial Analysis directly rather than taken on trust.
Watch next
- 01Independent, reproducible benchmark comparisons (not just Artificial Analysis) confirming the claimed 3.1% WER and 0.16s latency across diverse audio conditions, not just the tested 25 languages
- 02Developer adoption signals: usage volume, published case studies, or third-party integrations citing Muse Voice Transcribe within 3-6 months of launch
- 03Competitor pricing or capability responses from OpenAI, ElevenLabs, AssemblyAI, Deepgram and Cartesia in the following quarter, indicating whether this actually shifted the competitive floor
Coverage
1 storyMore capability signals
Full feed →- AI systems outperform expert humans in persuasive communication22 Jun 202681
- WIRED investigation: Flock Safety's AI person-search tools let police run broad description-based surveillance, with weak guardrails against misuse3 Sept 202680
- Hcompany open-sources NeoMME, a from-scratch multimodal-native encoder family, and NeoMME-Retriever for visual document retrieval3 Sept 202679
- Benchmark results show significant improvement in AI agent performance on WorkBench15 Jun 202679