Microsoft Launches New MAI Models for Real-Time Voice AI
Microsoft has expanded its artificial intelligence portfolio with the launch of MAI-Transcribe-2-Streaming, a real-time speech-to-text model, alongside two new voice-generation models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. The new models are designed to help developers build faster, more natural and multilingual conversational AI experiences.
MAI-Transcribe-2-Streaming Enables Real-Time Transcription
The new MAI-Transcribe-2-Streaming model is built for applications that need to process speech as it happens. Microsoft says the model can deliver low-latency transcription across 60 languages and automatically detect the language being spoken.
Instead of waiting for a speaker to complete a sentence, the model begins generating partial transcription results shortly after receiving the audio. Microsoft says its first hypotheses can appear in just over 100 milliseconds, after which the system continuously updates the text as additional speech context becomes available.
This approach can support voice assistants, customer service applications, live captions, dictation and other voice-based interfaces. Microsoft also highlights its potential...
Copyright of this story solely belongs to www.itvoice.in. To see the full text click HERE