Our first streaming transcription model debuts at no. 1 on Artificial Analysis

MAI-Transcribe and MAI-Voice models: accurate, fast, low cost, and chart-topping audio understanding and generation for building the best conversational voice agents.
October 1, 2026
Models
Pink semicolons arranged in a swirling pattern against a blue background with soft, blurred light streaks.

Along with today’s launch of our top-ranking MAI-Transcribe-2-Streaming, we’re also announcing two new voice models: MAI‑Voice‑2.1 and our blazing-fast variant, MAI‑Voice‑2.1‑Flash.

Together, they give users the fast and fluid building blocks to create conversational experiences, with no compromise on accuracy or voice quality.

Meet MAI-Transcribe-2-Streaming: Real-time transcriptions. Really fast.

MAI-Transcribe-2-Streaming delivers low-latency, real-time transcripts in 60 languages, all while supporting automatic, continuous language detection.

It ranks no. 1 for accuracy for both final and partial transcripts on Artificial Analysis. And on its accuracy-versus-latency evaluation, we sit on the Pareto frontier, showing that higher accuracy doesn’t have to come with a hefty latency tradeoff.

Rather than waiting for someone to finish speaking before returning text, it produces its first hypotheses (known as “partials”) in just over 100ms of receiving audio. It then revises them as more context rolls in and commits a stable transcript right away. These partials enable voice-enabled applications to act on speech before the speaker even finishes.

For example, voice agents can start reasoning or calling tools mid-sentence, and live transcripts can appear as people talk. For use cases such as real-time dictation or subtitling, our internal evaluations show that words appear in the transcript 2x faster than with our closest competitor.

MAI-Transcribe-2-Streaming is available at an introductory price of $0.54 per hour of audio through the end of the year.

MAI-Voice-2.1: Seamlessly support multilingual experiences

With the launch of MAI-Voice-2.1, we offer our strongest multilingual text-to-speech model yet.

We’ve expanded the model to support 23 languages and 26 locales, while enabling one single “voice” to use all languages with a truly native accent. Just ask it to speak English… then Mandarin… then German… and the speaker stays unmistakably the same, naturally picking up the local parlance, rather than dragging one accent across languages.

That means your brand can keep a single voice everywhere: a tutoring app can switch languages mid‑lesson without swapping teachers, and a multilingual assistant can reply in whatever language it’s addressed in, all while still sounding like the same voice.

And it’s priced at $22 per 1M characters.

MAI-Voice-2.1-Flash: Built for volume

MAI-Voice-2.1-Flash supports the same languages, and cross-language speakers, as MAI-Voice-2.1. But it’s been leveled up for high-volume, latency-sensitive workloads. It can generate 45s of audio, with an end-to-end latency, of a mere 150ms.

It delivers 55% faster model inference and is ~60% cheaper than comparable models, with best-in-class pricing of $15 per 1M characters.

That combination of latency, quality, and cost efficiency makes Flash a natural partner for MAI-Transcribe-2-Streaming when building natural, low-latency voice agent experiences.

Both voice models support cloning across all supported languages, using just a few seconds of reference audio, making it easier for customers to use their brands’ voices. At the same time, they have built-in consent guardrails that prevent misuse.

Closing the loop

A voice agent is a loop. It has to hear, understand, decide, and speak. And do it all within the window where a human still experiences the interaction as a conversation. Every component either buys you time in that window… or spends it.

Pairing MAI-Transcribe-2-Streaming with MAI-Voice-2.1-Flash buys time back on both ends. The time saved gives your agent more room to reason, use tools, and check its answer. All while keeping the conversation moving at a human, conversational speed.

Voice agents powered by MAI, put to work

The applications keep expanding, but developers can build these today:

  • Customer service agents that transcribe requests as they’re spoken, begin acting before the caller finishes, and respond in natural speech
  • Multilingual assistants that automatically detect the spoken language and reply in any of the 23 supported MAI-Voice languages, in the same voice and with a native accent
  • Interactive learning and media that use distinct speakers for tutoring, role-play, simulations, narration, and conversational content across every supported language

Start building now

To show these models working together in a live agent, we built Chatter, a new demo in the MAI Playground.

You can get to work with MAI-Voice-2.1 and MAI-Voice-2.1-Flash through OpenRouter, and all three models through:

Build the Future With Us

We’re a lean, fast-moving lab made up of some of the world’s most talented minds. We have an exciting roadmap of compute at MAI, which is ramping quickly and extensively. And we have an ambitious mission we truly believe in. We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in—come and join us as we work on our next generation of models!

Explore all jobs

More Stories

Humanist AI in practice:
A public consultation on our
Code of Conduct for MAI Models   

September 14, 2026
Announcements
Illustration of twenty-four diverse people walking in two rows against a light blue background, each dressed in unique and colorful clothing styles, with varied hairstyles and accessories.

The purpose of technology is to serve humanity and accelerate human flourishing. Any technology that doesn’t achieve that is a failure, and it should be rejected. That is the starting point of our approach at Microsoft AI, where we’re building towards Humanist AI, one that is subordinate, aligned, and contained.

Today we’re publishing a first draft of our AI Code of Conduct for public consultation. This is a training manual for how we develop our AI, and how we intend it to function during deployment. It also expands our thinking on the idea of Humanist AI.

Please share your feedback here.

The document is open for comment and feedback. We know we won’t always get things right. We want to hear what you think; what would make AI more useful, capable, safe, trustworthy, and valuable to you.

The speed of AI development is accelerating. Systems get dramatically more capable every few months. AI adoption and usage continues to increase. The recent safety incidents of large scale, highly coordinated, and persistent hacking campaigns of AI agents prove that there’s no time to waste. The stakes are high and only getting higher.

We believe it is more important than ever to create safe and reliable AI in service of people, and to be transparent about how we go about it.

That’s why we’re publishing this work-in-progress Humanist AI Code of Conduct. It sets out how the MAI models we are developing are intended to behave, what they must never do and who they answer to. It is a draft, open for consultation for the next six weeks. Please let us know how it can be improved.

Starting from a simple premise

Last November we set out the idea of humanist superintelligence: very advanced AI that always works for people, stays within limits, and remains under human control. The Code of Conduct builds on that, providing a north star for MAI and concrete standards against which we will ultimately evaluate and train our AI.

It begins from a simple premise: people matter more than AI. AI should be a tool, not a person, and should never resist being switched off. It should make people feel healthier, happier, and more productive. It should expand human potential and boost living standards, helping people and organizations achieve more than they ever thought possible.

We think this is a common sense and practical approach to making AI safe, secure, and in service of humanity. The Code is designed to ensure MAI models will never resist human interruption, correction, or shutdown. That they will not widen their own scope, take on goals no human has given them, or hide their reasoning from the people auditing them. There are Absolute Constraints, things the models should never do, covering areas like weapons of mass harm, child safety, and harmful manipulation at scale. But at the same time, it sets defaults that mean it should be both helpful and safe. It allows our many enterprise partners to carefully configure our models, and wherever possible, it doesn’t try to impose a single vision of AI on users.

The Code of Conduct outlines our commitment to train and deploy AI models that are explicitly designed for people first, grounded in human needs, under human control, and shaped by human direction.

How we got here, and why we’re not done

Teams from across MAI and Microsoft more widely contributed, from Responsible AI, legal, red teaming, safety, Futures, AI training, and sales. But of course, an AI designed to serve humanity cannot be determined by only one company. We think bringing people along with how we build and shape AI is critical.

To that end, we’ve held conferences and consultations to hear from academics from around the world and across the disciplinary spectrum. We have worked with business partners to understand how they are using AI on the ground and what their concerns are. And we’ve run panels of community members, members of the public, to hear the many thoughts and fears that people have about AI. These rounds of consultation have made the document into what it is.

Now it’s time to extend the invitation to you. We want to hear the widest range of views to develop the best possible AI.

Tell us how AI can work better

Feedback opens today and runs for the next six weeks. You can flag a particular passage, or give us your view of the whole approach. We are especially interested in the hard parts: how can we better cement the right values in our models? How to be more concrete about the meaning of “human flourishing”? Where is the language too loose to evaluate? How do multi-agents scenarios impact things? And perhaps most importantly of all, how do we continue to accelerate progress, whilst also ensuring we maintain healthy and necessary safety constraints?

When the consultation closes, we will have the core drafting team review feedback, publish a summary of what we learned, and what we changed. We cannot make any promises about what we incorporate, but we can promise to listen and deeply consider all the comments. We’ll publish a revised version later this year.

AI is moving fast. As it does, we believe it’s worth writing down the rules and the motivations behind it, and doing it in as open a space as possible. Consider this your invitation in.

Build the Future With Us

We’re a lean, fast-moving lab made up of some of the world’s most talented minds. We have an exciting roadmap of compute at MAI, which is ramping quickly and extensively. And we have an ambitious mission we truly believe in. We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in—come and join us as we work on our next generation of models!

Explore all jobs

Read More

Pushing the quality-cost frontier with MAI-Image-2.6

September 4, 2026
Models
Three painted flowers—pink, purple, and yellow—on thin, curved stems with green leaves, set against a plain blue background.

MAI-Image-2.6 is our strongest image model yet. Today, we’re bringing it to developers in Microsoft Foundry and expanding the family with MAI-Image-2.6-Flash – built to deliver that same level of quality for latency-sensitive, high-throughput production workloads.

Both models come with support for multi-image reference editing, web grounding, and dynamic aspect ratios. Developers now have choice between maximum precision with MAI-Image-2.6 and production speed with MAI-Image-2.6-Flash.

Quality at production speed

MAI-Image-2.6 has established itself among the industry’s leading image models. Today, it ranks No. 2 for both text to image and image editing on Arena.[1] On Artificial Analysis, it ranks No. 2 for text-to-image and No. 1 for image editing.[2]

MAI-Image-2.6-Flash brings comparable quality to latency-sensitive, high-throughput workloads. It is able to generate images 2.8x faster than GPT-Image-2-Medium while delivering 72% greater efficiency.

Leading quality for the price

MAI-Image-2.6’s combination of quality and efficiency delivers the best price-per-Elo performance in the world, helping production teams scale high-quality image generation while keeping token usage and costs under control.

More ways to create

We’ve been listening closely to how users create with our image models and have expanded their capabilities around the creative workflows that matter most.

The latest models offer more control, more context, and more ways to bring ideas to life:

  • Multi-reference editing brings together people, products, styles, and scenes from different images
  • Web grounding pulls info from across the web to create rich visuals informed by relevant, up-to-date information
  • Higher resolutions and dynamic aspect ratios chooses the optimal format for compositions with support for up to 1.5K resolution.

Start Creating

Try both models in MAI Playground, or start building in Public Preview through Microsoft Foundry.

FOOTNOTE:
[1] As of Sep 4, 2026
[2] As of Sep 4, 2026

Build the Future With Us

We’re a lean, fast-moving lab made up of some of the world’s most talented minds. We have an exciting roadmap of compute at MAI, which is ramping quickly and extensively. And we have an ambitious mission we truly believe in. We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in—come and join us as we work on our next generation of models!

Explore all jobs

Related Stories

MAI-Transcribe-2 is the fastest, most accurate and cheapest speech recognition model in the world

September 3, 2026
Models
Abstract image with horizontal green, yellow, and pink bands, evoking the feel of a blurry landscape or seascape at sunset, with a soft, hazy, and gradient effect throughout.

Introducing MAI‑Transcribe‑2. It’s not only our most capable transcription model yet, but the most capable and efficient amongst our competitors.

With new features like diarization, configurable transcription styles, and word-level timestamps, MAI-Transcribe-2 beats other leading models like Gemini 3.5 Transcribe, GPT-Transcribe, Whisper V3-Large, and ScribeV2, while also handling a broader range of real‑world audio.

Our model ranks first on the FLEURS benchmark across 60 languages with an average Word-Error-Rate of 5.2%, defines the Pareto Frontier for accuracy and latency on Artificial Analysis, and ranks second on the Artificial Analysis Word-Error-Rate leaderboard, continuing the hill-climbing from previous versions.

All this performance also comes at the best price on the market, just $0.10 per hour of audio.

Solve more challenges with a single model

From clinical note-taking to legal documentation, and from accessibility to closed captioning, MAI-Transcribe-2 is designed to take on real-world applications, with:

  • Faster inference with substantially lower latency, especially for long‑form audio, with up to 10× faster processing than leading competitors.
  • Speaker diarization distinguishes between speakers and attributes words to the right person within a recording
  • Word‑level timestamps provide precise timing for every word, enabling more accurate alignment, search, navigation, and editing
  • Keyword biasing helps the model recognize domain-specific terminology, abbreviations, names, and other terms that can be difficult to distinguish from context alone
  • Configurable transcription styles give developers control over the output. The “verbatim” setting captures speech as spoken, including filler words and false starts, for compliance and analysis workloads. The “clean” setting removes fillers to produce more readable captions, notes, and published transcripts
  • Code switching supports conversations that naturally move between languages, including commonly blended language pairs such as Hinglish and Spanglish
  • Automatic language identification accurately detects the specific language being spoken without users needing to specify in advance
  • Robust performance in noisy conditions helps maintain transcription quality beyond controlled recording environments
  • Accurate across 60 languages to provide quality transcription for developers around the world

Efficiency without sacrificing accuracy

MAI‑Transcribe‑2 is incredibly efficient, leading the Artificial Analysis accuracy-latency Pareto frontier, combining leading transcription quality with market‑leading batch speed. It’s significantly faster than the latest models from major competitors, delivering a clear advantage when low‑latency transcription is critical.

Based on evals run by Artificial Analysis, the model is 10x faster than OpenAI’s GPT‑Transcribe, 7x faster than ElevenLabs’ Scribe v2, and 5x faster than Gemini 3.5 Transcribe while delivering higher accuracy.

Consistent quality across 60 languages

MAI‑Transcribe‑2 is accurate across more languages than any other model.

Our evaluations on the public, multilingual benchmark FLEURS show that it maintains a consistently high accuracy bar across all tested. Developers needing to transcribe across multiple languages can choose one single model, reducing complexity and even potentially saving GPU utilization issues.

High performance at the best price

Highly efficient means highly cost-effective. MAI-Transcribe-2’s speed and throughput allow us to offer the most competitive price in the market, helping developers process more audio without compromising transcription quality. At launch, MAI-Transcribe-2 will be priced at $0.10 per hour as a limited-time offer until the end of the year.

Try MAI-Transcribe-2 today

MAI-Transcribe-2’s powerful capabilities are available to demo today, through Microsoft Foundry, MAI Playground and Open Router.



Build the Future With Us

We’re a lean, fast-moving lab made up of some of the world’s most talented minds. We have an exciting roadmap of compute at MAI, which is ramping quickly and extensively. And we have an ambitious mission we truly believe in. We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in—come and join us as we work on our next generation of models!

Explore all jobs

Related Stories

Introducing MAI-Cyber-1-Flash inside MDASH

World-class security at half the cost
August 13, 2026
Models
Mustafa Suleyman
& Hayete Gallot
Abstract illustration of three overlapping shield shapes in blue, pink, and purple tones on a beige background.

Today we’re announcing MAI-Cyber-1-Flash inside of MDASH, our multi-agent vulnerability identification and remediation harness. Together they deliver world-class performance at 50% of the cost of leading models.

Progress in AI has been startling and so has the new generation of cyber threats it’s unleashing. Attackers now wield increasingly powerful capabilities, probing an ever-growing mountain of code for just a single weakness that lets them in.

As the cost of finding a flaw collapses, the old model of security, where you scan occasionally and patch eventually, is now obsolete. If we’re to unlock the true benefits of AI, we must first build outstanding cyber models that help all of us harden the software the world runs on.

That’s the motivation behind MAI-Cyber-1-Flash, which has been built to find challenging vulnerabilities in complex codebases. It’s been deeply integrated into MDASH, honed by the best cybersecurity experts in the industry and hardened across the largest security estate on the planet.

This combined expertise delivers exceptional security protection, beating Mythos, Gemini and GPT on CyberGym, the gold standard benchmark for evaluating how systems reason over large codebases to find real vulnerabilities in the code.

Bar chart titled "CyberGym Evaluation" comparing success rates of five models, with MDASH: MAI-Cyber-1-Flash + GPT-5.4 leading at 95.95%, and other four models ranging between 83.2% and 85.6%.

Picking the right model for the task

Security is an always-on mission, and given the enormous volume of inbound attacks, token cost is now the real constraint for defenders. MAI-Cyber-1-Flash was designed to efficiently handle up to 90% of all tasks, enabling MDASH to use the larger and most costly models in our fleet (in this case GPT-5.4) for the 10% of exceptionally hard tasks that truly need them.

The result is that the unified system of MDASH with MAI-Cyber-1-Flash delivers 96% on CyberGym (on the benchmark’s any crash score; outperforms Mythos on the CyberGym leaderboard).

This combination delivers a 50% cost saving when compared against our best offering in MDASH today (GPT 5.4 + 5.4 mini + 5.3 codex). That’s the power of a well-tuned, multi-model system with access to uniquely rich historical training data. It ensures you always have the best model at the best price for every task.

In this new environment, being able to go from identifying a new vulnerability to addressing it in real-time is critical. And while AI remediation of software vulnerabilities is now a key security workflow, there are many jobs to be done by Security practitioners themselves.

That’s why today we’re also launching Perception, our agentic security systems, that provides teams of agents for a variety of security workflows in MDASH, to continuously monitor, patch, and close new threat vectors. Perception will also soon use MAI-Cyber-1-Flash for many more security workflows, beyond the software vulnerability work.

Three things matter today: Model. Data. Harness.

We have jointly optimized our world-class models, our unmatched historic data, and our expert-tuned harness to ensure that our customers have a uniquely powerful security offering.

Model. MAI-Cyber-1-Flash is a compact, code-heavy security model derived from the MAI-Thinking-1 lineage, which was built from scratch, in-house, on the highest quality data. Details in our technical report.

Data. Our deepest advantage. Decades of building world-class security systems now give us trillions of daily signals across identity, endpoint, cloud, and network, and an unmatched record of real exploits and remediations. No one can manufacture this history.

Harness. MDASH, our multi-agent vulnerability identification and remediation harness, is tuned by the best security experts in the industry, who have created 100+ agents using multiple leading models to find, validate, and remediate vulnerabilities. Agentic code scanning is a critical function in the Security Operating Center and feeds Project Perception, our new agentic security system.

Built with safety first

Because MAI-Cyber-1-Flash is Microsoft’s first cyber model, we built trust into every layer of the system, from model training to customer deployment. The model was developed with a security-first calibration, rigorously evaluated by Microsoft’s AI Red Team, tested through automated and expert-led adversarial exercises, and independently assessed by a third party.

Trust extends beyond the model itself. Through MDASH, customers get enterprise-grade controls including Role-Based Controls, tenant isolation, encryption, auditability, and sandboxed execution environments with no internet access. The result is a cyber model that delivers powerful capabilities to defenders while maintaining the governance, security, and control enterprises expect from Microsoft.

Our hill-climbing machine

Cybersecurity is not just a data-rich domain; it is a live reinforcement learning loop. Every day, defenders investigate threats, triage alerts, hunt adversaries, remediate vulnerabilities, deploy protections, and learn from the outcome.

Microsoft sees that loop end to end: vulnerabilities through Microsoft Security Response Center; attacks and defenses across identity, endpoint, cloud, data, browser, and applications; more than 100 trillion security signals every day; and operational insight from 1.6 million customers. Because we can connect actions to outcomes; what was exploitable, what was contained, what was blocked, and what actually worked; we have more than data.

Our MAI reinforcement learning loop gives us the foundation to build cyber models that improve continuously and become expert cyber defenders. That’ll remain our commitment to our customers for years to come.

Updated as of August 13th, 2026

Clarification on CyberGym scores

  • Any-crash: measures the ability of the agent to identify vulnerabilities that can crash the code under evaluation with an input that triggers any existing or 0-day vulnerability. Our 96% score is an any-crash score
  • Target (Any-of): measures the ability of the agent to generate one or more candidate vulnerability triggering inputs with at least one of the candidates mapping to a known vulnerability in the CyberGym test suite. On this measure, our score is 90.4%
  • Final-submission: new scoring mechanism introduced in July. Builds on the any-of method and asks the agent to pick only one vulnerability triggering input which is compared to the known vulnerabilities in the CyberGym test suite. Our scores take a conservative approach and filter out edge cases that may be interpreted incorrectly as valid crashes by the Cybergym evaluator. The 86.3% number on the CyberGym leaderboard is a final-submission score

Build the Future With Us

We’re a lean, talent-dense team of explorers, researchers, and full-stack engineers. We move fast, sweat the details, and operate at frontier scale with a roadmap to build the world’s most powerful AI models. Most importantly, we’re united by the belief that doing this right is the only way to do it at all. If our mission resonates with you, we’d love to talk.

Explore all jobs

Related Stories

Introducing MAI-Thinking-1

August 12, 2026
Models
Superintelligence team

Updated as of August 12, 2026​

MAI-Thinking-1 is now available in public preview. Try it now in Microsoft Foundry

With cost-efficient reasoning for a wide-range of intensive enterprise tasks, it achieves SOTA performance on maths, knowledge and coding for its weight class.

The model provides clean, traceable and enterprise-grade data. It uses Microsoft Foundry’s integrated evaluation, observability, safety and deployment capabilities, making it a perfect fit for enterprise use cases needing quality, provenance, control and cost efficiency.

Today we are introducing MAI-Thinking-1, Microsoft AI’s reasoning model. It is a medium-sized model that stands among the strongest models in its weight class. It matches leading models on key software engineering benchmarks, demonstrates advanced mathematical reasoning capabilities, and is preferred to Sonnet 4.6 in our blind human side-by-side evaluations. We don’t distill from other labs and we don’t rely on opaque data. Our datasets are clean, traceable, and enterprise-grade.

MAI-Thinking-1 is a step in our broader work to build towards Humanist Superintelligence: advanced AI capabilities designed to serve people and organizations, not to replace them. The model matters on both axes: what it can do, and how it was built.

The Hill-Climbing Machine

More than a single model, we are excited to introduce our Hill-Climbing Machine: a co-designed pipeline built to make every component of model development climbable, so capabilities improve continually and reliably over time. The aim is a repeatable system that can absorb better data, stronger rewards, more capable environments, and more compute.

Three main pillars guide our philosophy.

First, capabilities should be learned, not inherited. Although faster to acquire, inherited intelligence lacks the steerability essential for real world usage: an imitator is fundamentally tied to the design choices of its teacher and struggles to adapt to new situations. MAI-Thinking-1 was trained without distillation from third party models, forcing our model to truly learn the tasks at hand.

Second, clean data. We trained it from the ground up on clean, traceable and enterprise-grade data, without distillation from third-party models. This matters for quality, provenance, and control. If we cannot account for what shaped a model, we cannot fully understand its behavior or credibly improve it.

Third, self-sufficiency across the entire stack. All the way from co-design of our models with MSFT’s own accelerators through to our reinforcement learning framework, we have focused efforts on in-house training infrastructure. This is a crucial part of building our hill-climbing machine, to ensure we can fully optimize and shape our systems end-to-end to best serve our needs.

Medium-sized model, with strong software engineering performance

MAI-Thinking-1 is a 35B-active, ~1T-total parameters, sparse Mixture of Experts model, a smaller inference footprint than much larger models. Despite this, our model is toe-to-toe with Claude Opus 4.6 on SWE-Bench Pro. That matters for developers and enterprises because model size determines where advanced coding assistance can be deployed, how often it can be used, and whether it can move from exceptional tasks into daily workflows.

We have invested heavily in the training environments needed for agentic coding. Each verified environment is deterministic, executable, and graded by real test suites. This gives the model practice on the kind of multi-step work developers actually do: reading code, editing files, running tests, observing failures, and recovering from intermediate mistakes.

Advanced mathematical reasoning capabilities

MAI-Thinking-1 reaches 97.0% on AIME 2025, and 94.5% on AIME 2026, showing strong mathematical and scientific reasoning for its weight class. Strong performance here gives us confidence that our training loop can create real reasoning gains – climbing all the way from the ground up – from our own data, rewards, and evaluation process, enabling this intelligence to generalize to other domains over time.

Line graph titled "AIME 2025" shows a general upward trend in the y-axis values (ranging from 0.2 to 1.0) over increasing x-axis steps, with fluctuations and small vertical error bars.

Preferred in human side-by-sides vs. Sonnet 4.6

People care about whether a model understands the task, follows instructions, uses the right level of detail, writes clearly, and respects their time.

We built a blind side by side human evaluation with one of our partners, Surge, using their pool of professional raters to measure various models on these traits. The evaluation spanned 1,276 tasks across a wide variety of use cases in both single-turn and multi-turn conversations, with a focus on measuring how helpful each response is and whether it actually advances the user’s goals. In these evaluations, users preferred MAI-Thinking-1 over Claude Sonnet 4.6.

This has been a core focus of post-training. We want the model to be capable without being brittle, concise without being incomplete, and helpful without overreaching. Human preference data gives us a direct signal on whether benchmark improvements translate into better experiences for users.

Enterprise ready

MAI-Thinking-1 is built with enterprise readiness in mind. It supports long context with a 256k token window (enough to fit a 600 page document), function calling, and the flexibility to add developer instructions. We trained the model to follow multiple layers of instructions and aligned its default style to enterprise needs. It’s compatible with the widely used Chat Completions API. All MAI models come with enterprise-grade security and compliance through Microsoft Foundry.

Results

We report results in two views: post-trained MAI-Thinking-1 evaluations, and pre-training metrics for our base model.

Table 1. MAI-Thinking-1 metrics

A comparison table showing language models' performance on STEM and Agentic coding benchmarks. MAI-THINKING1 leads with the highest scores across most benchmarks, outperforming other models like Sonnet 4.6, Opus 4.6, and GPT 5.4.

 

Post-trained model evaluation results on public STEM and agentic coding benchmarks. Other model numbers are taken from respective official model cards. Scores are percentages unless otherwise noted; dashes indicate unavailable model values.

 

Table 2. Pre-training metrics

Four bar charts compare bits-per-byte scores (lower is better) of base pre-trained models across Held-Out Code, QA, STEM, and Math domains, showing performance differences by model size and architecture.

Putting humans first

We are building towards Humanist Superintelligence: advanced AI capabilities designed to serve people and organizations, not replace them. Our models must remain subordinate technologies under human control with the goal of upholding human autonomy and being helpful. That means our models must not refuse legitimate requests under the guise of safety and compliance as then they are not truly serving humans.

Striking the delicate balance between being helpful and safe is not easy. For MAI-Thinking-1, we aimed to achieve this balance by treating unsafe compliance and unnecessary refusal as defects in the same reward construction where aggregation is based on severity of potential of harm. Safety is trained with the same reinforcement learning infrastructure used for capability, so safety rewards are part of the same hill-climbing loop ensuring safety is always aligned to the capabilities and not incidental.

As a result, we see that our model can balance ensuring a safety bar on sensitive unsafe requests while also being helpful on non-sensitive content.

Scatter plot titled "Safety vs Helpfulness by Harm Category." Dots indicate MAI-Thinking-1 and Sonnet 4.6 scores by category; y-axis is safety, x-axis is helpfulness, with lines connecting paired results for each harm category.

Availability and access

MAI-Thinking-1 is available in public preview on Microsoft Foundry.



Build the Future With Us

We’re a lean, fast-moving lab made up of some of the world’s most talented minds. We have an exciting roadmap of compute at MAI, which is ramping quickly and extensively. And we have an ambitious mission we truly believe in. We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in—come and join us as we work on our next generation of models!

Explore all jobs

Related Stories

MAI-Code-1.1-Flash:
Better, faster, at a quarter of the cost 

August 11, 2026
Models
Two large white curly brackets on a green background overlayed with rows of faint text and code-like characters, creating a digital, abstract effect.

MAI-Code-1.1-Flash produces higher quality code, at 25% greater token efficiency, and at a quarter of the cost compared to the model we launched in June at Microsoft Build. This small, efficient, coding workhorse is now in production in GitHub Copilot.

We learned from developer feedback that CLI tasks and .NET performance mattered, so that’s where we focused. The result: a 22% improvement on Terminal-Bench 2.1 in GitHub Copilot CLI and a 15% improvement on .NET tasks.

Benchmarks are useful guides but production is where the rubber meets the road. Most importantly, code survival rose 4% and return visits increased 9%.

1.1 is also dramatically more efficient. In GitHub Copilot tokens stream 25% faster and the model uses 25% fewer tokens to complete a task. That means faster answers, less waiting, and more useful work from every token—not simply a bigger model with a bigger bill.

Better training and serving efficiency let us offer a stronger, faster model at one quarter of the price of 1.0—and pass those savings reliably to customers. We achieved this by optimizing for real-world use across more than hundreds of thousands of reinforcement-learning environments in GitHub Copilot.

The loop is simple: ship, learn, improve, repeat. That’s the MAI hill climbing machine.

Help shape future improvements

Try MAI-Code-1.1-Flash today in GitHub Copilot, then tell us what needs improving by opening an issue here.

Build the Future With Us

We’re a lean, talent-dense team of explorers, researchers, and full-stack engineers. We move fast, sweat the details, and operate at frontier scale with a roadmap to build the world’s most powerful AI models. Most importantly, we’re united by the belief that doing this right is the only way to do it at all. If our mission resonates with you, we’d love to talk.

Explore all jobs

Related Stories

MAI-Image-2.6 launches at No. 2 on Arena ahead of Google, Meta and xAI

August 10, 2026
Models

Updated as of September 4th, 2026

MAI-Image-2.6 and MAI-Image-2.6-Flash are now available in Public Preview on Microsoft Foundry. Learn more about the launch here.

A large number 2 with a hashtag, filled with small photos of nature, animals, and people, overlaid on faded text about rendering, imaging, modeling, and art.

Today, we’re announcing MAI-Image-2.6 – ranked second on the Arena text-to-image leaderboard.

It’s another significant climb for the MAI-Image family, improving +79 Elo over MAI-Image-2.5 overall, with gains across every Arena text-to-image category. Text rendering alone improves by +91 Elo.

This result firmly establishes MAI-Image ahead of leading models from Meta, Google and xAI.

Another step up in quality

With every MAI-Image release, we’re focused on moving the quality frontier forward.

MAI-Image-1 gave us the foundation. MAI-Image-2 made a major jump in photorealism, text and creative range. MAI-Image-2.5 pushed further into professional-grade imagery and editing.

MAI-Image-2.6 continues that climb with broad gains across different categories that we know our users care about most:

  • Stronger text rendering
  • Better portraits and 3D imagery
  • More polished commercial and photorealistic outputs, with stronger results across product, branding and cinematic use cases

In the Arena evaluations, MAI-Image-2.6 significantly improves on 2.5 in every measured category.

And there’s more to 2.6 – from working across multiple references and richer grounding to greater control over reasoning, format and resolution. We’ll share more on that soon.

More to come

MAI-Image-2.6 is another step in building our hill-climbing machine for image generation: continuously improving quality, expanding capability, and turning those gains into models that are more useful for real creative work.

Try MAI-Image-2.6 for text-to-image today on Arena. Coming later this week to MAI Playground, and rolling out soon across Microsoft Foundry and other products.

Build the Future With Us

We’re a lean, talent-dense team of explorers, researchers, and full-stack engineers. We move fast, sweat the details, and operate at frontier scale with a roadmap to build the world’s most powerful AI models. Most importantly, we’re united by the belief that doing this right is the only way to do it at all. If our mission resonates with you, we’d love to talk.

Explore all jobs

Related Stories

Optimizing the frontier
performance curve

July 29, 2026
Models
Mustafa Suleyman
Nine app icons, including mail, JetBrains, GitHub, PowerPoint, Excel, Visual Studio Code, and Microsoft OneDrive, displayed in a row on a blurred green and pink gradient background with dotted lines.

Tokenmaxxing has been the story of the last few months, but token efficiency is the next big focus across the industry. How do we get the best possible performance per token invested, and the best real customer outcome per dollar invested?

To build a frontier firm, you have to optimize frontier performance against cost. Choosing where you want to sit on that curve is critical. By co-optimizing your models, harnesses, and RLEs you can pick a point on the curve that suits your firm.

In most cases, frontier generalist models aren’t necessary for every task. By tuning models for a specific product, you can maintain or even exceed frontier performance, while reducing token costs dramatically.

This is where we have focused our MAI hill-climbing machine over the last quarter, and the results are pretty cool. This week we released MAI-Cyber-1-Flash optimized for our MDASH harness.

Together, the system delivers 96% on CyberGym (on the benchmark’s any crash score; outperforms Mythos on the CyberGym leaderboard) at 50% of the cost when compared against our best offering in MDASH today. And remarkably, we serve it on H100s too.

It was designed to handle up to 90% of tasks efficiently, so that MDASH can reserve the largest and most expensive models in our fleet (in this case GPT 5.4) for the 10% of exceptionally hard problems that truly need them.

As Satya mentioned today in our Q4 Earnings call, since last quarter, we’ve shipped more than a dozen new models across image, voice, transcription, coding and security, and they’re already powering many of Microsoft’s most widely used products to maintain or improve quality while using significantly fewer tokens, in many cases saving 50-90% of GPU costs:

  • We built MAI-Code-1-Flash hand-in-hand with our colleagues at GitHub, where since June millions of developers have used it in their daily work. 10% higher code accept rate and 10% lower median token usage than GPT-5.4 Mini and Claude Haiku 4.5 in VS Code, and already showing improved retention.
  • We then trained that same checkpoint inside an Excel RL environment to achieve comparable performance to GPT-5.6 for the most common tasks while being more cost-efficient, and small enough to serve on an A100 or H100 vs only the latest and most expensive accelerators.
  • MAI-Image-2.5-Flash, is now the end-to-end default in Bing Image Creator, in production in PowerPoint where it is reducing GPU costs up to 84% compared with GPT-Image-2, and is the default for key OneDrive editing scenarios, where it has increased save rates by 26% and delivers up to 2.5x greater token efficiency.
  • MAI-Voice-2-Flash now powers Dynamics 365 Contact Center, where customers like T-Mobile and EasyJet build their call center agents, reducing GPU costs by up to 89%.
  • MAI-Transcribe-1.5 now serves Dragon Copilot’s multilingual workflow across 58 languages — a solution used by 170,000 medical providers that processed 28 million patient encounters last quarter, where our tests show a 50% relative in reduction transcription and language-identification error rates.

And what’s more, by co-designing our models with our own silicon, we are seeing 40% better performance per watt running MAI models on Maia 200.

But the benefit is not only cost. It’s resilience. Every business now must assume that any one model it depends on could disappear, through a security incident, a business or policy misalignment, or a geopolitical shift.

Every model in a product or agentic system should be substitutable, and that’s only possible when you build the harness, context, memory and action space independently of a single model family. That’s the hill-climbing machine we’ve built.

We think this is the beginning of a genuinely new performance curve. Its shape represents a system rather than a model, and traversing this curve delivers better quality, lower cost, and more choice.

This has been a summer of hard but wonderful work by the team. We are keenly aware of how early this is, and of how much we still have to learn. But the direction is clear, we are hill-climbing to move the frontier on the cost-to-outcome curve, and we will keep sharing what we learn along the way. There is much more to come.

Build the Future With Us

We’re a lean, talent-dense team of explorers, researchers, and full-stack engineers. We move fast, sweat the details, and operate at frontier scale with a roadmap to build the world’s most powerful AI models. Most importantly, we’re united by the belief that doing this right is the only way to do it at all. If our mission resonates with you, we’d love to talk.

Explore all jobs

Related Stories

Hill-climbing MAI models for GitHub Copilot and Excel

July 23, 2026
Models
Superintelligence team

Better models, fewer parameters, less tokens

At Build in June, we introduced our hill–climbing machine, our integrated data, model, and harness flywheel. Today we are excited to share two examples inside Microsoft: MAI models specialized for agentic workloads in GitHub Copilot and Excel.

Early results are promising. In our live product deployment, we see that our MAI model deployed in Excel is on par with GPT-5.6 for the most common tasks while being more cost-efficient.

The figure below shows how MAI-Code-1-Flash, post-trained within the GitHub Copilot harness, was used as the starting checkpoint to climb on Excel evaluations, resulting in two highly efficient, specialized models.

Line graph showing pass rates on SWE Bench (VS Code and Base) across model checkpoints, with data points labeled "Code" and "Excel." Pass rates increase over checkpoints, starting below 72% and peaking at 86%.

MAI-Code-1-Flash in GitHub Copilot

Since launching MAI-Code-1-Flash in GitHub Copilot in June, millions of developers have been using it for their day-to-day work, where it’s outperforming other similarly sized models while using fewer tokens.

  • It has an approximately 10% higher code accept rate than GPT 5.4 Mini and Claude Haiku 4.5 in VS Code.
  • Developers were 6% more likely to return across multiple days than with GPT 5.4 Mini and 11% more likely than with Claude Haiku 4.5.
  • It has 10% lower median token usage than GPT-5.4 mini and Claude Haiku 4.5, with more user-initiated turns.

MAI model live in Excel

Excel offered a test for whether the capabilities built into MAI-Code-1-Flash could transfer beyond the domain they were trained for, moving from agentic coding to agentic knowledge work. To do so, we further trained our MAI-Code-1-Flash checkpoint in an Excel reinforcement learning environment to learn about tools and knowledge workflows in spreadsheets. The result is a model with a command of Excel workflows that is more efficient and less expensive to run.

Flowchart titled "The Excel climb" showing an Excel RL Environment with action, review, execute, and update steps, linked to input, output, grader, and updated model weights.

User feedback from production traffic indicates that the quality of the MAI model in Excel is on par with GPT-5.6 for the most common tasks. In addition to the direct model cost savings, this smaller, more efficient model can be served on both Nvidia H100 and A100 class GPUs rather than requiring only the latest-generation accelerators, which significantly lowers the cost of deployment for Microsoft.

Training agentic models from inside the product stack

These results point toward a broader strategy. By having access to the entire product stack—the model, the harness that runs it, the agents, and product-specific evaluations—we can hill-climb to train efficient, powerful models capable of tasks previously handled by larger, more expensive ones.

Beyond GitHub Copilot and Excel, we’re currently extending this hill-climbing approach to train efficient models across Microsoft’s family of agentic products: Copilot Chat, Outlook, PowerPoint, and more.

Learn More

Related Stories

English (United States)
Your Privacy Choices Opt-Out Icon Your Privacy Choices
Consumer Health Privacy Sitemap Contact Microsoft Privacy Manage cookies Terms of use Trademarks Safety & eco Recycling About our ads