Our first streaming transcription model debuts at no. 1 on Artificial Analysis

MAI-Transcribe and MAI-Voice models: accurate, fast, low cost, and chart-topping audio understanding and generation for building the best conversational voice agents.
October 1, 2026
Models
Pink semicolons arranged in a swirling pattern against a blue background with soft, blurred light streaks.

Along with today’s launch of our top-ranking MAI-Transcribe-2-Streaming, we’re also announcing two new voice models: MAI‑Voice‑2.1 and our blazing-fast variant, MAI‑Voice‑2.1‑Flash.

Together, they give users the fast and fluid building blocks to create conversational experiences, with no compromise on accuracy or voice quality.

Meet MAI-Transcribe-2-Streaming: Real-time transcriptions. Really fast.

MAI-Transcribe-2-Streaming delivers low-latency, real-time transcripts in 60 languages, all while supporting automatic, continuous language detection.

It ranks no. 1 for accuracy for both final and partial transcripts on Artificial Analysis. And on its accuracy-versus-latency evaluation, we sit on the Pareto frontier, showing that higher accuracy doesn’t have to come with a hefty latency tradeoff.

Rather than waiting for someone to finish speaking before returning text, it produces its first hypotheses (known as “partials”) in just over 100ms of receiving audio. It then revises them as more context rolls in and commits a stable transcript right away. These partials enable voice-enabled applications to act on speech before the speaker even finishes.

For example, voice agents can start reasoning or calling tools mid-sentence, and live transcripts can appear as people talk. For use cases such as real-time dictation or subtitling, our internal evaluations show that words appear in the transcript 2x faster than with our closest competitor.

MAI-Transcribe-2-Streaming is available at an introductory price of $0.54 per hour of audio through the end of the year.

MAI-Voice-2.1: Seamlessly support multilingual experiences

With the launch of MAI-Voice-2.1, we offer our strongest multilingual text-to-speech model yet.

We’ve expanded the model to support 23 languages and 26 locales, while enabling one single “voice” to use all languages with a truly native accent. Just ask it to speak English… then Mandarin… then German… and the speaker stays unmistakably the same, naturally picking up the local parlance, rather than dragging one accent across languages.

That means your brand can keep a single voice everywhere: a tutoring app can switch languages mid‑lesson without swapping teachers, and a multilingual assistant can reply in whatever language it’s addressed in, all while still sounding like the same voice.

And it’s priced at $22 per 1M characters.

MAI-Voice-2.1-Flash: Built for volume

MAI-Voice-2.1-Flash supports the same languages, and cross-language speakers, as MAI-Voice-2.1. But it’s been leveled up for high-volume, latency-sensitive workloads. It can generate 45s of audio, with an end-to-end latency, of a mere 150ms.

It delivers 55% faster model inference and is ~60% cheaper than comparable models, with best-in-class pricing of $15 per 1M characters.

That combination of latency, quality, and cost efficiency makes Flash a natural partner for MAI-Transcribe-2-Streaming when building natural, low-latency voice agent experiences.

Both voice models support cloning across all supported languages, using just a few seconds of reference audio, making it easier for customers to use their brands’ voices. At the same time, they have built-in consent guardrails that prevent misuse.

Closing the loop

A voice agent is a loop. It has to hear, understand, decide, and speak. And do it all within the window where a human still experiences the interaction as a conversation. Every component either buys you time in that window… or spends it.

Pairing MAI-Transcribe-2-Streaming with MAI-Voice-2.1-Flash buys time back on both ends. The time saved gives your agent more room to reason, use tools, and check its answer. All while keeping the conversation moving at a human, conversational speed.

Voice agents powered by MAI, put to work

The applications keep expanding, but developers can build these today:

  • Customer service agents that transcribe requests as they’re spoken, begin acting before the caller finishes, and respond in natural speech
  • Multilingual assistants that automatically detect the spoken language and reply in any of the 23 supported MAI-Voice languages, in the same voice and with a native accent
  • Interactive learning and media that use distinct speakers for tutoring, role-play, simulations, narration, and conversational content across every supported language

Start building now

To show these models working together in a live agent, we built Chatter, a new demo in the MAI Playground.

You can get to work with MAI-Voice-2.1 and MAI-Voice-2.1-Flash through OpenRouter, and all three models through:

Build the Future With Us

We’re a lean, fast-moving lab made up of some of the world’s most talented minds. We have an exciting roadmap of compute at MAI, which is ramping quickly and extensively. And we have an ambitious mission we truly believe in. We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in—come and join us as we work on our next generation of models!

Explore all jobs

More Stories

English (United States)
Your Privacy Choices Opt-Out Icon Your Privacy Choices
Consumer Health Privacy Sitemap Contact Microsoft Privacy Manage cookies Terms of use Trademarks Safety & eco Recycling About our ads