On October 1, Microsoft announced MAI-Transcribe-2-Streaming, which returns transcription results while a person is still speaking, along with two speech synthesis models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. All three are available in public preview on Microsoft Foundry, so they are not yet at a stage where production use is guaranteed.

The release adds real-time speech recognition and speech synthesis to the recording-oriented transcription model announced in September. Processing can begin before the user finishes speaking, but the transcription returned midway can be rewritten later.

How close the new models bring voice agents to practical use depends on how far the tentative recognition results obtained mid-utterance are relied on in downstream processing.

AD

Treat mid-utterance recognition results as subject to rewriting

MAI-Transcribe-2-Streaming supports 60 languages and continuously auto-detects the language being spoken.

According to Microsoft, it returns the first tentative transcription a little over 100 milliseconds after receiving audio. Unlike methods that recognize speech in one batch after a recording ends, this allows processes such as displaying captions or pre-searching for information relevant to a question to proceed while the user is still talking.

However, the text returned midway does not necessarily become the beginning of the final result as is. When more audio arrives, earlier recognition results may be rewritten.

Microsoft's Realtime API documentation draws a clear distinction in how to handle this.

Characters delivered via delta are confirmed text and are appended to the end. Content delivered via the MAI-specific intermediate replaces the entire unconfirmed portion at that point.

If an app simply appends every received result in order, the text before and after a correction will end up duplicated.

"Confirmed" here only means confirmed as a transcription. It does not necessarily mean the user's intent is settled.

Suppose, for example, that in a booking conversation a user corrects themselves:

"Next Friday—no, Thursday."

The app can start searching availability using the interim recognition of "Friday." But if it finalizes the booking at that moment, it would be an error once the speech is corrected to "Thursday."

This is not a product test but an example of processing design using streaming recognition.

Separating processes that can be undone later, such as searches and candidate retrieval, from those that change actual state, such as confirming a booking or issuing a refund, lets developers use mid-utterance results while limiting erroneous actions.

Detecting utterance boundaries is also up to the app

The app must also supply its own logic for judging where an utterance ends.

The Realtime API does not support automatic server-side detection of end of speech or automatic commit requests. The app must judge natural pauses or the end of a recording and send commit itself.

When combining voice activity detection, how a short breath is distinguished from a genuine end of speech affects response speed.

Judging the boundary too early starts processing while the user is still speaking. Waiting too long fails to take full advantage of the low latency that streaming recognition offers.

The session limit is one hour.

The API sends and receives audio in a format similar to the OpenAI Realtime API. However, the MAI-specific intermediate and the conditions for sending a commit request are not identical.

Rather than assuming full compatibility because the API formats resemble each other, developers need to check whether interim results are updated correctly and whether conversation boundaries are handled properly.

AD

"WER 2.5%" and "just over 100 ms" are not the same kind of metric

In the Artificial Analysis evaluation that Microsoft listed in the model card for the Streaming version, the word error rate (WER) of the final transcription was 2.5%.

In the same table, Grok Transcribe 2.0 scored 2.7% and ElevenLabs Scribe V2 Realtime scored 3.6%.

WER is calculated by adding up the substitutions, insertions, and deletions of words relative to a reference text and dividing by the number of words in the reference. The lower the value, the fewer the transcription errors.

But the 2.5% cannot be taken as an average accuracy across all 60 languages, including Japanese.

Artificial Analysis's streaming evaluation uses about eight hours of audio, combining the conversational dataset AA-AgentTalk at 50% with VoxPopuli and Earnings22 at 25% each.

According to its methodology, the target language for VoxPopuli is English.

Therefore, "supports 60 languages" and "has been confirmed to deliver the same accuracy in all 60 languages" need to be considered separately.

For the standard MAI-Transcribe-2, announced on September 3, there is also a Microsoft measurement showing an average WER of 5.2% on FLEURS, which covers 60 languages.

However, the model and the evaluation data differ from the Streaming version's 2.5%.

Comparing the numbers alone does not support a conclusion that errors have been cut by more than half.

The starting point for measuring latency also differs by metric

Speed figures also need to be read for what they mean.

The "just over 100 milliseconds" that Microsoft describes refers to the time from receiving audio to returning the first tentative recognition result.

Artificial Analysis's latency metric, by contrast, starts measuring at the point when SileroVAD detects the end of speech, and includes network latency.

In other words,

  • how quickly the first characters arrive during speech
  • how quickly the result returns after the speaker finishes

are different metrics.

Comparing them simply as the same "latency" risks misjudging behavior in actual use.

The same applies to speech synthesis.

The MAI-Voice-2.1 comparison table lists model inference latency of about 550 milliseconds for the standard version and about 45 milliseconds for the Flash version.

Pricing per million characters is $22 for the standard version and $15 for Flash.

Microsoft also describes the Flash version as having "end-to-end latency 150ms" when generating 45 seconds of audio.

However, the published materials do not define which processes that 150 milliseconds includes in as much detail as they do for the 45-millisecond model inference latency.

In any case, simply adding the recognition figure of just over 100 milliseconds to the synthesis figure of 45 milliseconds does not give the time until the user hears a reply.

Between them come steps such as:

  • understanding the content of the question
  • generating an answer with a language model
  • search and database lookups
  • running external tools
  • processing until audio playback begins

The published latency values help in choosing individual components, but the response speed of a voice agent as a whole must be measured in the actual configuration.

AD

Real-time operation costs some features

Under the specifications Microsoft had published as of October 2, 2026, both the standard MAI-Transcribe-2 and the Streaming version support 60 languages.

The Streaming version, however, lacks some features the standard version has, such as speaker diarization and correction for specialized terms.

The table below organizes the official version comparison table by the same items.

Prices are in US dollars per hour of audio, comparing the introductory prices through the end of 2026 listed in the standard version's announcement and the Streaming version's model card.

Item MAI-Transcribe-2 (standard) MAI-Transcribe-2-Streaming
Supported languages 60 languages 60 languages
Streaming recognition during speech Not supported Supported
Speaker diarization Supported Not supported
Word-level timestamps Supported Not supported
Recognition correction via specified keywords Supported Not supported
Choice of cleaned-up or verbatim text Supported Not supported
Introductory price per hour of audio $0.10 $0.54

The standard version has features suited to keeping meetings and calls as records after the fact.

It can also distinguish who spoke and make it easier to recognize business-specific terms.

The Streaming version, by contrast, emphasizes keeping up with the user's speech as quickly as possible and passing the content to the next process in real time.

At present, in exchange for real-time operation, some features for producing records that are easy to reread later must be provided separately.

This does not rule out the possibility that these features will be added in the future.

Using different models during and after a call

Taking advantage of this difference, one could build a setup that recognizes and responds with the Streaming version during a call, then reprocesses with the standard version after the call ends to produce a record including speaker diarization and timestamps.

This is not a demonstrated result from an actual deployment but one example of how the published features could be combined.

Processing the same audio twice would naturally incur additional costs.

Still, it makes possible a design that uses different models for different purposes:

  • prioritize response speed during the conversation
  • prioritize record quality afterward

Comparing introductory prices alone, the Streaming version's per-hour audio rate is 5.4 times that of the standard version.

0.54 ÷ 0.10 = 5.4

However, this ratio does not include downstream costs such as language models and speech synthesis.

The Streaming version's $0.54 is also an introductory price through December 31, 2026, and the rate after that cannot be determined from these materials.

Estimating the cost per call requires accounting for downstream processing fees and when the introductory price ends.

23 languages with the same voice, but no Japanese speech synthesis

The languages supported by MAI-Voice-2.1 and the Flash version differ from the 60 languages of the speech recognition model.

The model card for the standard version and the model card for the Flash version list support for 23 languages.

The models are said to maintain a consistent speaker-like voice quality when switching between supported languages, and to let developers specify emotion and speaking style per utterance or sentence.

One could imagine uses such as a multilingual help desk that keeps the same character voice while changing how it speaks depending on the content.

Voice cloning uses 5 to 60 seconds of reference audio, and Microsoft says no additional training is required.

However, this does not mean anyone can freely replicate a voice with just a short recording.

In addition to Microsoft's approval, a consent recording made by the person providing the voice is required. The developer documentation also describes the application and consent procedures.

Because the Flash version can output generated audio progressively, it can also be used for "barge-in": stopping playback and moving on to the next step when a user starts speaking in the middle of the AI's speech.

But a fast synthesis model does not automatically yield natural barge-in handling.

At what point to stop playback, and how to reflect what the user restated in the next answer, must be designed across the whole app.

Japanese can be recognized but not spoken by the new synthesis models

For Japanese use, speech recognition and speech synthesis need to be considered separately.

Japanese is included in the supported-language list for MAI-Transcribe-2-Streaming.

On the other hand, Japanese does not appear in the supported-language lists for MAI-Voice-2.1 and MAI-Voice-2.1-Flash, announced this time.

They support English, Chinese, Korean, and others, but there is no basis for saying a Japanese voice conversation can be completed with these three models alone.

If Japanese recognition results are paired with a separate Japanese-capable speech synthesis model, voice quality and response time would need to be verified again in that configuration.

A gap remains between public preview and production use

Microsoft Learn positions both MAI-Transcribe-2-Streaming and the speech synthesis models as public preview.

There is no SLA guaranteeing service levels, and use in production environments is not recommended.

A stage at which a prototype can be built is not the same as one at which a customer service desk can be entrusted around the clock.

What has come together this time is:

  • processing that converts speech to text
  • processing that generates speech from text

The language model that understands the meaning of a question in between, looks up the necessary business information, and composes an answer must be chosen separately.

If the time taken by recognition and synthesis shrinks, that time could potentially be spent on verifying answers and using external tools.

On the other hand, if the search target or database is slow to respond, the conversation as a whole will reply slowly even if only the voice models are fast.

Practicality must be measured across the whole conversation

Before adoption, the whole conversation needs to be measured in the actual usage environment, not just through benchmarks of individual models.

Important points include, for example:

  • How many seconds it takes, on a real network connection, from when the user finishes speaking until audio playback begins
  • Whether self-corrections like "Friday—no, Thursday" are correctly reflected in processes such as bookings
  • Whether playback can be stopped immediately when the user interrupts the AI's speech
  • Whether the next answer after an interruption correctly uses the corrected content
  • How accurately Japanese personal names, product names, and business terminology are recognized

Recognition accuracy including Japanese business terminology cannot be judged from the rankings presented this time alone.

Microsoft has increased the options for starting the next process while a user is still speaking.

But whether this becomes practical is determined by more than the standalone performance of speech recognition and synthesis.

The requirements are that the system does not take wrong actions when users correct themselves, can continue the conversation naturally when interrupted, and can be run at a realistic cost even after the introductory price ends.

If these conditions can be confirmed across the whole conversation, the tentative recognition results obtained mid-utterance can lead to voice responses that users can rely on with confidence.