Meta launched its new speech-to-text model, ‘Muse Voice Transcribe’, on September 1. According to the company, the model can convert spoken words into text in real time during conversations and understand multiple languages ​​within a single interaction.
Muse Voice Transcribe is the first real-time audio perception model introduced by Meta Superintelligence Labs (MSL). The model was trained on more than 70 languages, with 25+ languages extensively validated at launch. These include Hindi and major Indian languages ​​such as Tamil, Telugu, Malayalam, and Kannada.
Meta says that ‘Muse Voice Transcribe’ has been developed specifically with scenarios like code-switching in mind. This means a person can easily switch between different languages ​​during a conversation, and the model can handle this transition without requiring additional models or separate processing. According to the company, this capability can make multilingual conversations more seamless while also improving the speed and utility of real-time transcription.
Muse Voice Transcribe could boost real-time transcription in India, where conversations often span English and regional languages. Meta says the model can transcribe speech as it happens, distinguish 20+ speakers, and handle recordings longer than an hour.
Meta has launched Muse Voice Transcribe through its Model API, offering speech-to-text at $3 per 1,000 audio minutes (about $0.18 per hour). The model is already powering dictation in Meta AI for Mac and Muse Code, and is designed to handle transcription, coding, and voice-assistant tasks, with support for code-switching and multiple regional languages that could make it especially useful in multilingual markets such as India.


