READ. SCROLL. LISTEN.

Original briefings. Zero spin.

Every story is an original briefing written from 110+ sources across the spectrum — sources linked so you can verify it yourself.

← Back to headlines

Meta Ships Muse Voice Transcribe, Claims Top Spot on a Benchmark It Cited Itself

Meta Ships Muse Voice Transcribe, Claims Top Spot on a Benchmark It Cited Itself
Meta Superintelligence Labs launched Muse Voice Transcribe on September 1, a real-time transcription model that handles up to 20+ speakers, 70-plus languages and mid-sentence code-switching in one system. Meta says it ranks first on the Artificial Analysis streaming speech-to-text leaderboard, but that's Meta's own citation and no independent lab has confirmed it yet.

Meta Superintelligence Labs, the division CEO Mark Zuckerberg built to chase artificial general intelligence, launched its first product on Tuesday, September 1. It's not a chatbot or a reasoning engine. It's a transcription model called Muse Voice Transcribe.

The model does three jobs most transcription tools handle separately: real-time speech-to-text, speaker diarization (figuring out who's talking), and endpointing (knowing when someone actually stopped talking versus just pausing). Meta's research blog says it does all three in a single autoregressive model, processing audio in 80-millisecond chunks and deciding, chunk by chunk, whether to keep listening or start writing text.

Zuckerberg posted a demo on X, his first return to the platform in three years of not posting there, showing the model switching between speakers and languages mid-conversation, including "code-switching," where a speaker blends two languages in the same sentence. Alexandr Wang, who leads product for the effort, posted that the model is "live now via meta model api" and already powering dictation in the Meta AI Mac app and Muse Code.

What it actually does

According to Meta's own research post, the model was trained across more than 70 languages, with 25 "extensively checked" for this launch. News9Live reported that Hindi, Tamil, Telugu, Kannada and Malayalam get native support, a detail also confirmed by IANS, which quoted Meta's statement directly. That's a real gap other Western AI labs have been slower to close, since code-switching between English and regional languages is common in markets like India.

The diarization piece can track more than 20 speakers across recordings longer than an hour, without a separate speaker-labeling step afterward, according to Meta and KuCoin's reporting. Pricing is $3 per 1,000 audio minutes, which News9Live calculated works out to about 18 cents an hour.

The benchmark claim needs a caveat

Meta says Muse Voice Transcribe ranks first on Artificial Analysis's streaming speech-to-text leaderboard and on public diarization benchmarks, "as of September 1, 2026," per its own research blog. News9Live cited Meta's chart showing a 3.1% error rate on the streaming test and a 17.5% error rate on diarization, ahead of AssemblyAI's U3.5 Pro Offline (21.1%), ElevenLabs Scribe v2 Offline (24.6%) and Deepgram Nova 3 Offline (25.4%).

Those numbers come from Meta's own presentation of a third-party benchmark, not from an independent test run by a competitor or a neutral outside party. News9Live itself flagged that "benchmark results do not always reflect every real-world accent, microphone or noisy room." KuCoin's coverage went further, noting "the absence of independent benchmarks is worth noting" and that without head-to-head testing on standard datasets, "it's hard to evaluate where Muse Voice Transcribe sits relative to existing options on raw accuracy."

Companies picking which benchmark to publish, and when, is standard practice across the AI industry, not unique to Meta. It doesn't mean the numbers are wrong. It means nobody outside Meta has independently reproduced them yet.

KuCoin's report also stated Meta "hasn't specified exactly which languages or how many" and that "specific pricing details... have not been publicly released." Both claims are contradicted by Meta's own research blog and by News9Live and IANS, which quote the 70-plus-languages figure, the 25-language validation number, and the $3-per-1,000-minutes price directly from Meta's statement. That looks like an editorial gap, not a fact in dispute.

The competitive backdrop

Engadget reported that Meta's launch came less than a week after Google shipped Gemini 3.5 Transcribe, its own real-time audio model, though Engadget noted it's unclear whether Meta plans to bake Muse Voice Transcribe into flagship apps like Instagram or WhatsApp the way Google is integrating its model into Android and eventually Chrome. For now, Meta's model lives in a narrower lane: the Meta AI Mac app, Muse Code, and the Meta Model API for developers.

The Muse Spark family the model belongs to is Meta Superintelligence Labs' third or fourth release in a matter of weeks, following a coding agent, an open-weight model, and the Mac app itself. Whether any of it turns into a product ordinary users touch daily, rather than a developer tool and a benchmark chart, is still an open question Meta hasn't answered.

Sources used for this briefing

This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.

center-left
EngadgetMeta's new AI transcription model can distinguish between multiple speakers and languages in real-time
unknown
Ground NewsMeta Just Beat OpenAI and Google at Real-Time Transcription
unknown
The New StackMeta just beat OpenAI and Google at real-time transcription
unknown
research.meta.aiIntroducing Muse Voice Transcribe
unknown
News9LiveMeta Muse voice transcribe launches with Hindi, Tamil, Telugu, Kannada and Malayalam support
unknown
KuCoinMeta's MSL Launches Muse Voice Transcribe, Real-Time Audio Model with Speaker Diarization
unknown
Ian's LiveMeta launches 'Muse Voice Transcribe' with native support for five Indian languages