Meta has launched Muse Voice Transcribe, a real-time speech recognition system designed to transcribe multilingual conversations while identifying individual speakers.
Developed by Meta Superintelligence Labs, the model combines automatic speech recognition with speaker diarisation and endpoint detection, allowing it to determine who is speaking and when individual turns begin and end.
Meta says the model was trained using more than 70 languages, with 25 extensively validated for its initial release. It can also recognise code-switching, where speakers move between languages within the same sentence or conversation.
Muse Voice Transcribe processes audio continuously and adjusts the amount of context it waits for before producing individual words, an approach Meta says is designed to balance transcription accuracy against latency. The system can handle audio exceeding an hour and conversations involving more than 20 speakers without requiring separate post-processing.
The technology is available through the Meta Model API and has also been integrated into Meta AI for Mac and Muse Code, extending Meta's push into voice-driven AI assistants and productivity tools.