Microsoft released three new voice AI models on October 1 local time, including the real-time speech-to-text model MAI-Transcribe-2-Streaming, and the text-to-speech models MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Among them, MAI-Transcribe-2-Streaming achieved the highest accuracy in streaming speech recognition tests on the AI benchmark platform Artificial Analysis, with a word error rate (WER) of 2.5% and a latency of approximately 0.13 seconds, surpassing SpaceXAI Grok Voice Transcribe 2.0. This model supports 60 languages and is available at a promotional price of $0.54 per hour until the end of the year. MAI-Voice-2.1 supports 23 languages, can maintain voice consistency across languages, and is priced at $22 per million characters; MAI-Voice-2.1-Flash targets cost-sensitive applications with lower latency and a price of $15 per million characters, with a 55% increase in inference speed.