top of page
20250531_095654.avif

Meta Muse Voice Transcribe Supports 5 Indian Languages in Real Time

The company Meta has come up with a new tool called Muse Voice Transcribe, which is an audio perception model that works in real-time and has more than 70 language models, including Hindi, Tamil, Telugu, Kannada, and Malayalam.

Meta Muse Voice Transcribe

Meta has released a new speech recognition model, Muse Voice Transcribe, aimed at providing fast and high-accuracy real-time transcription of audio files. This model is capable of supporting more than 70 languages around the world and is natively compatible with five major languages spoken in India Hindi, Tamil, Telugu, Kannada and Malayalam. According to Meta, this is the first-ever real-time audio perception model created by their Meta Superintelligence Labs with the integration of multiple audio processing features into one unit.


In contrast to traditional transcription services which first analyze an entire audio file before turning it into a text file, the Muse Voice Transcribe will be able to transcribe speech in real time while recognizing speakers and detecting the beginning and the end of their speech. According to Meta, these abilities all work in parallel, so there is no necessity of going through another post-processing phase.

Real-Time Transcription With Multi-Speaker Support

Another great feature of Muse Voice Transcribe is that it can transcribe long and complicated talks. As Meta states, the model can identify more than 20 speakers within one recording, which makes it suitable for meetings, interviews, group discussions, and any other conversation between different people. It can also transcribe audio recordings longer than an hour. In this case, developers do not have to cut the recording into several pieces because the model can deal with it.


The way Muse Voice Transcribe works is that it processes audio data in very small chunks, which last 80 milliseconds. It does not commit every single word to text at once but rather estimates how many more milliseconds it requires for the most precise transcription. If some words are easy to recognize, the result will be given immediately, but complicated words will get more processing time.


It is especially crucial in real-time transcription applications where the trade-off between the two is inevitable since making a transcription faster could mean making mistakes and taking longer time could mean the transcription will not feel coherent in the discussion. Muse Voice Transcribe was created to provide an effective balance between both by dynamically determining when there is enough evidence for each word.

Muse Voice Transcribe Supports Five Major Indian Languages

In terms of users and developers from India, one of the most distinctive features of this new model is that it supports five popular Indian languages. The model is able to process the following five languages without any issues: Hindi, Tamil, Telugu, Kannada and Malayalam. Alongside with the support of the above mentioned languages, Muse Voice Transcribe also supports a number of other languages.


According to Meta, in the course of creating the model they used information in more than 70 languages, and validated 25 of those languages before releasing the first version. Language diversity could be useful for developers who want to create transcription, accessibility, productivity and voice applications for international users.


It also worth mentioning the ability of this new model to process code-switching in conversations. Code-switching is the switching from one language to another during the same conversation. Muse Voice Transcribe can detect such cases automatically without asking users to change language every time when someone changes the language. Even in the middle of the phrase.


In addition, Muse Voice Transcribe has options to apply biasing based on languages, keywords, and context. Through these features, the service will have access not only to the data available in the audio but also to the context in which the conversation is being held to make a more precise choice when recognizing certain words. It might be helpful for recognizing people’s names, technical vocabulary, brand names, and other words that might confuse the speech recognition system.


Thus, through integrating real-time transcription, speaker recognition, endpointing, multilingual support, and contextual analysis, Meta’s product has much more functionality compared to regular automatic speech-to-text software. While standard models work just with an audio file, this one works as a real-time audio perception system that understands the structure of the ongoing conversation.

Meta Muse Voice Transcribe Availability and Pricing

Muse Voice Transcribe has been released by Meta to developers through the Meta Model API, making it possible for developers to add the speech transcription model in third-party applications. The organization has stated that even their applications such as Meta AI for Mac and Muse Code, use the same technology for transcribing dictations.


The cost for using the Meta Model API is $3 per 1,000 audio minutes, which comes down to about $0.18 per hour or about Rs. 17 per hour according to the provided conversion. This simple price structure may prove to be quite interesting for those developers who require real-time speech transcription without having to develop their own speech recognition system.


By offering support for more than 70 languages, five major languages of India, long duration recordings and conversations of more than 20 participants, Muse Voice Transcribe is a major milestone towards Meta's goal of developing real-time AI audio processing capabilities. Being capable of handling multilingual and code-switched conversations, the new tool may prove to be quite useful for a market like India, where people switch languages mid-conversation.


Subscribe to our newsletter

Comments


bottom of page