Meta Muse Voice Transcribe: Live Dictation Powered by AI

Meta Muse Voice Transcribe: Live Dictation Powered by AI

Meta Muse Voice Transcribe types along live and tells speakers apart in real time – free as a Mac dictation feature, paid per API call for developers.

Too much jargon?→ Look it up in the glossary

You talk, and the text shows up while you're still mid-sentence. No "send" button, no waiting for the AI to start only after your last word. That's exactly what Meta's new model Muse Voice Transcribe promises – and for the first time in years, dictating on a computer feels new again, instead of like the clunky Windows speech recognition of the 2000s.

What Muse Voice Transcribe actually does

Muse Voice Transcribe comes out of Meta Superintelligence Labs and handles three jobs in a single pass: turning speech into text (transcription), figuring out who's talking (diarization), and noticing when someone's done talking (endpointing). Every 80 milliseconds, the model decides anew whether to keep listening or start writing. Sounds like engineering trivia, but it's the difference between "AI types with a noticeable lag" and "AI types along almost simultaneously."

Meta trained the model on more than 70 languages, naming concrete quality numbers for around 25 of them – without specifying exactly which ones. In practice, that means: expect solid results for English, German, or Spanish, and test the more exotic languages yourself before trusting them. The model also handles code-switching – jumping between two languages mid-sentence, the way real conversations actually work.

The speaker recognition is the impressive part: the model claims to tell apart up to 20 people at once. For meeting notes or a podcast transcript, that's a big deal – assuming real-world use lives up to the benchmarks.

What's free – and what actually costs money?

This is where it gets interesting for two pretty different audiences. If you own a Mac, you basically get Muse Voice Transcribe for free already: it's baked into the dictation feature in Meta AI for Mac. Hold the Fn key, talk, let go – the text lands in practically any app, no extra download required.

If you want to build something with it yourself – a voice bot, live captions, your own transcription tool – you pay through the Meta Model API, currently around $3 per 1,000 audio-minutes. That works out to roughly 18 cents per hour of audio. Not nothing, but not a number that ruins a hobby developer either.

The catch, as usual: there's no way to self-host the model. If you need everything running on-premise for privacy reasons, or you're already committed to open weights, Muse Voice Transcribe leaves you out in the cold. For everyone else, the API is convenient – but convenient also means your audio data travels through Meta's servers.

Privacy: where do your words end up?

Meta states that for its own demo, audio is processed only to produce the transcript and isn't stored afterward – and it has users confirm beforehand that they have permission to record. What that means in practice for third-party apps built on the API, though, is anyone's guess: the model itself makes no blanket promise covering every use case. Before you let something record an important one-on-one conversation live, it's worth a quick check of what that particular app actually does with the data – not because anyone's assumed to be doing something shady, but because you can't rule it out either.

Who is this actually for?

The Mac dictation feature is worth it for anyone who types a lot and is tired of the click-through hassle of dictation apps – no setup, works right away, and that makes it genuinely useful for beginners too. The API, on the other hand, is developer territory: if you're not writing code yourself, you won't notice much beyond Muse Voice Transcribe quietly showing up in apps you already use over the coming months.

Either way, it's worth a look: Mac users can try the Fn key next time they're typing a chat message, and everyone else can just keep an eye out for where Muse Voice Transcribe turns up next. Dictation hasn't been this quietly good in a while.