• About Us
King of Computer Media
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
King of Computer Media
No Result
View All Result

Home - AI Trends and Related News - Meta Launches Muse Voice Transcribe, a Real-Time Speech-to-Text Model Supporting Over 70 Languages and Distinguishing Between More Than 20 Speakers

Meta Launches Muse Voice Transcribe, a Real-Time Speech-to-Text Model Supporting Over 70 Languages and Distinguishing Between More Than 20 Speakers

Rocky by Rocky
September 2, 2026
in AI Trends and Related News

As AI technology keeps advancing, speech-to-text models can now work in real time—meaning text appears as you speak. Meta recently launched the new Muse Voice Transcribe real-time speech-to-text model, which boasts a word error rate of just 3.1% in streaming (outputting words as you speak) mode, and takes only 0.16 seconds from when you finish speaking to when the final text is produced. It can even identify which of more than 20 speakers is talking throughout. It’s currently available for free in the Meta AI Mac version, and the official website also offers a free trial feature.

Meta releases Muse Voice Transcribe real-time speech-to-text model: 3.1% streaming word error rate ranks first, can distinguish more than 20 speakers, and pressing and holding Fn on Mac enables system-wide voice typing.

Earlier, Meta Superintelligence Labs (MSL) officially launched Muse Voice Transcribe., this is MSL’s first real-time audio perception model, built on the Muse Spark model family that Meta released in April, and is an autoregressive multimodal model.

The biggest difference from typical speech-to-text services is that what previously required three systems working in tandem can now be done by a single model, including real-time streaming speech recognition, speaker identification (capable of distinguishing more than 20 speakers in the same audio segment), and endpointing to determine “whether this person has finished speaking.” The training covers over 70 languages, 25 of which were fully validated at launch, and it supports switching languages mid-sentence or between sentences. It can also transcribe long audio files over an hour long with more than 20 speakers, without requiring any post-processing.

Real-time speech-to-text has a major challenge: if the model outputs text too early, it may make mistakes because it hasn’t heard enough information; but if it always waits a long time before outputting, it loses the meaning of “real-time.”

Muse Voice Transcribe’s approach is to let the model itself decide how long to “listen” to each word. The audio is split into 80-millisecond chunks, or 12.5 chunks per second. Each time a chunk is received, the model determines whether the current information is sufficient to output text. If not, it waits for the next chunk; if it is, it begins outputting text.

Meta calls this mechanism “adaptive delay.” The wait time for different words is not fixed; for words that are harder to judge, the model can listen a bit longer before deciding, while for words that are easier to recognize, it can output faster, balancing recognition accuracy and response speed.

On the training side, Meta uses reinforcement learning, combining the word error rate (WER) reward with a latency reward in a multiplicative manner, so the model cannot optimize for speed or accuracy alone, but learns to strike a balance between the two.

According to the Artificial Analysis speed-versus-accuracy scatter plot shared by Meta, Muse Voice Transcribe’s final transcription word error rate is only 3.1%, and it takes only about 0.16 seconds from when the system detects that the user has finished speaking to when the final text is produced, placing it below the Pareto frontier (the optimal boundary for speed and accuracy). In short, at the same speed, no other model is more accurate than it; at the same accuracy, no other model is faster than it:

Compared with competitors, Soniox v5 Real-Time takes only about 0.05 seconds, but its word error rate is 4.5%; ElevenLabs Scribe v2 Realtime is similar in speed to Muse (about 0.15 seconds), but its word error rate is also 3.6%; Google’s Gemini 3.5 Transcribe Live, released last month, comes in at 0.41 seconds and 4.0%; OpenAI’s GPT Live Transcribe takes as long as 0.84 seconds, with a word error rate of 3.9%.

Below is the complete ranking of final transcription word error rates. For this English test, a 3.1% WER can be roughly understood as about 3.1 transcription errors per 100 English words; the lower the number, the better:


For speaker identification, the traditional common approach is to split speech-to-text and speaker clustering into separate processing pipelines, with some systems even requiring post-processing after the audio finishes. Muse Voice Transcribe, by contrast, integrates diarization directly into the same streaming model—when the model detects a speaker change, it inserts <|start_of_turn|> and the corresponding speaker tag into the output, so no separate speaker diarization step is needed.

Meta also shared test data for Diarization speaker identification error rates, using the average of three public datasets: AMI-IHM, AMI-SDM, and VoxConverse. Muse Voice Transcribe scored 17.5%. More notably, Muse Voice Transcribe achieved this in a real-time streaming state, already significantly outperforming the offline results of AssemblyAI, ElevenLabs, and Deepgram.

However, Meta didn’t compare with Google’s Gemini 3.5 Transcribe in this chart, nor with Soniox, the fastest one—it feels like Meta specifically cherry-picked.

Meta said that when training this model, it covered more than 70 languages, of which 25 were fully validated at launch, including Chinese, Japanese, French, Spanish, Hindi, Vietnamese, and others. However, the official list of 25 languages was not fully disclosed. In my testing, it seems Traditional Chinese is not yet supported; the output is all Simplified Chinese.

As for where you can experience Muse Voice Transcribe, there are currently three ways:

  • The speech-to-text used by Meta AI for Mac is now this model, completely free to use, with global support.
  • Muse Code
  • Meta Model API pricing: $3 per 1,000 audio minutes ($0.18 per hour)

After opening Meta AI for Mac, use the voice-to-text feature to experience this new model:

Currently, no matter how I phrase it, the output is always in Simplified Chinese.

Voice-to-text related settings can be found in the Dictation menu, such as: enabling hotkeys, permissions, input devices, etc.:

Source: KOCPC Chinese

Tags: aiMACMETAMuse Voice TranscribeSpeech-to-text

Recent Posts

  • Gemini can now help you remember where you’ve stored your items.
  • Meta Launches Muse Voice Transcribe, a Real-Time Speech-to-Text Model Supporting Over 70 Languages and Distinguishing Between More Than 20 Speakers
  • Google launches ‘Google Pics’, a new AI-powered generative image creation and editing tool.
  • NVIDIA unveils NVHBM: Not making new memory, but seizing the discourse power from Samsung and SK Hynix.
  • Dyson launches the Dyson CameraJet™ micro-jet electric toothbrush! Water flossing, brushing, and rinsing all in one device, priced at 16,990 yuan.

Recent Comments

No comments to show.
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology

No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology