About This Article
This article was generated using an automated workflow powered by generative AI. Based on official Microsoft AI and Microsoft Learn information from October 1, 2026, it outlines the availability and key implementation considerations for MAI-Transcribe-2-Streaming. No API execution was performed.Verification Status: 📘 Verified via official Microsoft primary sources. API execution unverified.
Information Verification Date: October 2, 2026.
Microsoft AI has released MAI-Transcribe-2-Streaming. This model returns intermediate hypothesis partial transcripts before the speaker finishes talking, updating them to final results dynamically.
What Is New
Microsoft states that the model supports 60 languages and returns the first partial transcript in just over 100ms from audio ingestion. Microsoft Learn outlines a setup using WebSockets to continuously stream audio and handle intermediate, delta, and completed events.
However, It Is in Public Preview
Microsoft Learn explicitly notes that the feature is currently in Public Preview, carries no SLA, and is not recommended for production workloads. Each session has a maximum duration of 1 hour. Usage requires a Microsoft Foundry resource and a model deployment.
Available regions include Sweden Central, Central US, and South India, while East US 2 is listed as coming soon. Even with global access, verifying the processing region is necessary.
Checklist for Administrators and Developers
Do not deploy previews directly to production environments
Establish audio data retention and transfer policies
Manage Microsoft Entra bearer tokens or API keys securely
Ensure application logic does not confuse partial and final transcripts
Monitor region availability, quotas, and session limits
Evaluate recognition error rates alongside latency using production data
Pricing
Microsoft AI has announced an introductory price of $0.54 per audio hour until the end of 2026. Future pricing must be verified separately.

