MAI-Transcribe-2-Streaming Public Preview: Real-Time Speech Transcription Before Utterance Completion

Microsoft 365・Azureカテゴリを表すパンダのイラスト Microsoft 365 / Azure

About This Article
This article was generated using an automated workflow powered by generative AI. Based on official Microsoft AI and Microsoft Learn information from October 1, 2026, it outlines the availability and key implementation considerations for MAI-Transcribe-2-Streaming. No API execution was performed.

Verification Status: 📘 Verified via official Microsoft primary sources. API execution unverified.
Information Verification Date: October 2, 2026.

Microsoft AI has released MAI-Transcribe-2-Streaming. This model returns intermediate hypothesis partial transcripts before the speaker finishes talking, updating them to final results dynamically.

What Is New

Microsoft states that the model supports 60 languages and returns the first partial transcript in just over 100ms from audio ingestion. Microsoft Learn outlines a setup using WebSockets to continuously stream audio and handle intermediate, delta, and completed events.

However, It Is in Public Preview

Microsoft Learn explicitly notes that the feature is currently in Public Preview, carries no SLA, and is not recommended for production workloads. Each session has a maximum duration of 1 hour. Usage requires a Microsoft Foundry resource and a model deployment.

Available regions include Sweden Central, Central US, and South India, while East US 2 is listed as coming soon. Even with global access, verifying the processing region is necessary.

Checklist for Administrators and Developers

  • Do not deploy previews directly to production environments

  • Establish audio data retention and transfer policies

  • Manage Microsoft Entra bearer tokens or API keys securely

  • Ensure application logic does not confuse partial and final transcripts

  • Monitor region availability, quotas, and session limits

  • Evaluate recognition error rates alongside latency using production data

Pricing

Microsoft AI has announced an introductory price of $0.54 per audio hour until the end of 2026. Future pricing must be verified separately.

Official Information and Primary Sources

Document information

Article title
MAI-Transcribe-2-Streaming Public Preview: Real-Time Speech Transcription Before Utterance Completion
Published
Updated
Source
https://papanda925.com/?p=17699&lang=en

License: Text and original figures for which this site holds the relevant rights are available under CC BY 4.0 , unless otherwise noted. This article may include content created or edited with generative AI. If code has a separate license notice or a linked GitHub repository license, that license takes precedence for the code. Quotations, third-party materials, images, and trademarks are excluded from this license. Usage policy

Copied title and URL