Our first streaming transcription model debuts at no. 1 on Artificial Analysis
Along with today’s launch of our top-ranking MAI-Transcribe-2-Streaming, we’re also announcing two new voice models: MAI‑Voice‑2.1 and our blazing-fast variant, MAI‑Voice‑2.1‑Flash.
Together, they give users the fast and fluid building blocks to create conversational experiences, with no compromise on accuracy or voice quality.
Meet MAI-Transcribe-2-Streaming: Real-time transcriptions. Really fast.
MAI-Transcribe-2-Streaming delivers low-latency, real-time transcripts in 60 languages, all while supporting automatic, continuous language detection.
It ranks no. 1 for accuracy for both final and partial transcripts on Artificial Analysis. And on its accuracy-versus-latency evaluation, we sit on the Pareto frontier, showing that higher accuracy doesn’t have to come with a hefty latency tradeoff.
Rather than waiting for someone to finish speaking before returning text, it produces its first hypotheses (known as “partials”) in just over 100ms of receiving audio. It then revises them as more context rolls in and commits a stable transcript right away. These partials enable voice-enabled applications to act on speech before the speaker even finishes.
For example, voice agents can start reasoning or calling tools mid-sentence, and live transcripts can appear as people talk. For use cases such as real-time dictation or subtitling, our internal evaluations show that words appear in the transcript 2x faster than with our closest competitor.
MAI-Transcribe-2-Streaming is available at an introductory price of $0.54 per hour of audio through the end of the year.
MAI-Voice-2.1: Seamlessly support multilingual experiences
With the launch of MAI-Voice-2.1, we offer our strongest multilingual text-to-speech model yet.
We’ve expanded the model to support 23 languages and 26 locales, while enabling one single “voice” to use all languages with a truly native accent. Just ask it to speak English… then Mandarin… then German… and the speaker stays unmistakably the same, naturally picking up the local parlance, rather than dragging one accent across languages.
That means your brand can keep a single voice everywhere: a tutoring app can switch languages mid‑lesson without swapping teachers, and a multilingual assistant can reply in whatever language it’s addressed in, all while still sounding like the same voice.
And it’s priced at $22 per 1M characters.
MAI-Voice-2.1-Flash: Built for volume
MAI-Voice-2.1-Flash supports the same languages, and cross-language speakers, as MAI-Voice-2.1. But it’s been leveled up for high-volume, latency-sensitive workloads. It can generate 45s of audio, with an end-to-end latency, of a mere 150ms.
It delivers 55% faster model inference and is ~60% cheaper than comparable models, with best-in-class pricing of $15 per 1M characters.
That combination of latency, quality, and cost efficiency makes Flash a natural partner for MAI-Transcribe-2-Streaming when building natural, low-latency voice agent experiences.
Both voice models support cloning across all supported languages, using just a few seconds of reference audio, making it easier for customers to use their brands’ voices. At the same time, they have built-in consent guardrails that prevent misuse.
Closing the loop
A voice agent is a loop. It has to hear, understand, decide, and speak. And do it all within the window where a human still experiences the interaction as a conversation. Every component either buys you time in that window… or spends it.
Pairing MAI-Transcribe-2-Streaming with MAI-Voice-2.1-Flash buys time back on both ends. The time saved gives your agent more room to reason, use tools, and check its answer. All while keeping the conversation moving at a human, conversational speed.
Voice agents powered by MAI, put to work
The applications keep expanding, but developers can build these today:
- Customer service agents that transcribe requests as they’re spoken, begin acting before the caller finishes, and respond in natural speech
- Multilingual assistants that automatically detect the spoken language and reply in any of the 23 supported MAI-Voice languages, in the same voice and with a native accent
- Interactive learning and media that use distinct speakers for tutoring, role-play, simulations, narration, and conversational content across every supported language
Start building now
To show these models working together in a live agent, we built Chatter, a new demo in the MAI Playground.
You can get to work with MAI-Voice-2.1 and MAI-Voice-2.1-Flash through OpenRouter, and all three models through:
- Microsoft Foundry
- MAI Playground
- Vercel
- LiveKit (coming soon)
- Azure Voice Live
Build the Future With Us
We’re a lean, fast-moving lab made up of some of the world’s most talented minds. We have an exciting roadmap of compute at MAI, which is ramping quickly and extensively. And we have an ambitious mission we truly believe in. We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in—come and join us as we work on our next generation of models!
More Stories
Humanist AI in practice:
A public consultation on our
Code of Conduct for MAI Models
A public consultation on our
Code of Conduct for MAI Models
The purpose of technology is to serve humanity and accelerate human flourishing. Any technology that doesn’t achieve that is a failure, and it should be rejected. That is the starting point of our approach at Microsoft AI, where we’re building towards Humanist AI, one that is subordinate, aligned, and contained.
Today we’re publishing a first draft of our AI Code of Conduct for public consultation. This is a training manual for how we develop our AI, and how we intend it to function during deployment. It also expands our thinking on the idea of Humanist AI.
Please share your feedback here.
The document is open for comment and feedback. We know we won’t always get things right. We want to hear what you think; what would make AI more useful, capable, safe, trustworthy, and valuable to you.
The speed of AI development is accelerating. Systems get dramatically more capable every few months. AI adoption and usage continues to increase. The recent safety incidents of large scale, highly coordinated, and persistent hacking campaigns of AI agents prove that there’s no time to waste. The stakes are high and only getting higher.
We believe it is more important than ever to create safe and reliable AI in service of people, and to be transparent about how we go about it.
That’s why we’re publishing this work-in-progress Humanist AI Code of Conduct. It sets out how the MAI models we are developing are intended to behave, what they must never do and who they answer to. It is a draft, open for consultation for the next six weeks. Please let us know how it can be improved.
Starting from a simple premise
Last November we set out the idea of humanist superintelligence: very advanced AI that always works for people, stays within limits, and remains under human control. The Code of Conduct builds on that, providing a north star for MAI and concrete standards against which we will ultimately evaluate and train our AI.
It begins from a simple premise: people matter more than AI. AI should be a tool, not a person, and should never resist being switched off. It should make people feel healthier, happier, and more productive. It should expand human potential and boost living standards, helping people and organizations achieve more than they ever thought possible.
We think this is a common sense and practical approach to making AI safe, secure, and in service of humanity. The Code is designed to ensure MAI models will never resist human interruption, correction, or shutdown. That they will not widen their own scope, take on goals no human has given them, or hide their reasoning from the people auditing them. There are Absolute Constraints, things the models should never do, covering areas like weapons of mass harm, child safety, and harmful manipulation at scale. But at the same time, it sets defaults that mean it should be both helpful and safe. It allows our many enterprise partners to carefully configure our models, and wherever possible, it doesn’t try to impose a single vision of AI on users.
The Code of Conduct outlines our commitment to train and deploy AI models that are explicitly designed for people first, grounded in human needs, under human control, and shaped by human direction.
How we got here, and why we’re not done
Teams from across MAI and Microsoft more widely contributed, from Responsible AI, legal, red teaming, safety, Futures, AI training, and sales. But of course, an AI designed to serve humanity cannot be determined by only one company. We think bringing people along with how we build and shape AI is critical.
To that end, we’ve held conferences and consultations to hear from academics from around the world and across the disciplinary spectrum. We have worked with business partners to understand how they are using AI on the ground and what their concerns are. And we’ve run panels of community members, members of the public, to hear the many thoughts and fears that people have about AI. These rounds of consultation have made the document into what it is.
Now it’s time to extend the invitation to you. We want to hear the widest range of views to develop the best possible AI.
Tell us how AI can work better
Feedback opens today and runs for the next six weeks. You can flag a particular passage, or give us your view of the whole approach. We are especially interested in the hard parts: how can we better cement the right values in our models? How to be more concrete about the meaning of “human flourishing”? Where is the language too loose to evaluate? How do multi-agents scenarios impact things? And perhaps most importantly of all, how do we continue to accelerate progress, whilst also ensuring we maintain healthy and necessary safety constraints?
When the consultation closes, we will have the core drafting team review feedback, publish a summary of what we learned, and what we changed. We cannot make any promises about what we incorporate, but we can promise to listen and deeply consider all the comments. We’ll publish a revised version later this year.
AI is moving fast. As it does, we believe it’s worth writing down the rules and the motivations behind it, and doing it in as open a space as possible. Consider this your invitation in.
Build the Future With Us
We’re a lean, fast-moving lab made up of some of the world’s most talented minds. We have an exciting roadmap of compute at MAI, which is ramping quickly and extensively. And we have an ambitious mission we truly believe in. We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in—come and join us as we work on our next generation of models!
Read More
Pushing the quality-cost frontier with MAI-Image-2.6
MAI-Image-2.6 is our strongest image model yet. Today, we’re bringing it to developers in Microsoft Foundry and expanding the family with MAI-Image-2.6-Flash – built to deliver that same level of quality for latency-sensitive, high-throughput production workloads.
Both models come with support for multi-image reference editing, web grounding, and dynamic aspect ratios. Developers now have choice between maximum precision with MAI-Image-2.6 and production speed with MAI-Image-2.6-Flash.
Quality at production speed
MAI-Image-2.6 has established itself among the industry’s leading image models. Today, it ranks No. 2 for both text to image and image editing on Arena.[1] On Artificial Analysis, it ranks No. 2 for text-to-image and No. 1 for image editing.[2]
MAI-Image-2.6-Flash brings comparable quality to latency-sensitive, high-throughput workloads. It is able to generate images 2.8x faster than GPT-Image-2-Medium while delivering 72% greater efficiency.
Leading quality for the price
MAI-Image-2.6’s combination of quality and efficiency delivers the best price-per-Elo performance in the world, helping production teams scale high-quality image generation while keeping token usage and costs under control.
More ways to create
We’ve been listening closely to how users create with our image models and have expanded their capabilities around the creative workflows that matter most.
The latest models offer more control, more context, and more ways to bring ideas to life:
- Multi-reference editing brings together people, products, styles, and scenes from different images
- Web grounding pulls info from across the web to create rich visuals informed by relevant, up-to-date information
- Higher resolutions and dynamic aspect ratios chooses the optimal format for compositions with support for up to 1.5K resolution.
Start Creating
Try both models in MAI Playground, or start building in Public Preview through Microsoft Foundry.
FOOTNOTE:
[1] As of Sep 4, 2026
[2] As of Sep 4, 2026
Build the Future With Us
We’re a lean, fast-moving lab made up of some of the world’s most talented minds. We have an exciting roadmap of compute at MAI, which is ramping quickly and extensively. And we have an ambitious mission we truly believe in. We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in—come and join us as we work on our next generation of models!
Related Stories
MAI-Transcribe-2 is the fastest, most accurate and cheapest speech recognition model in the world
Introducing MAI‑Transcribe‑2. It’s not only our most capable transcription model yet, but the most capable and efficient amongst our competitors.
With new features like diarization, configurable transcription styles, and word-level timestamps, MAI-Transcribe-2 beats other leading models like Gemini 3.5 Transcribe, GPT-Transcribe, Whisper V3-Large, and ScribeV2, while also handling a broader range of real‑world audio.
Our model ranks first on the FLEURS benchmark across 60 languages with an average Word-Error-Rate of 5.2%, defines the Pareto Frontier for accuracy and latency on Artificial Analysis, and ranks second on the Artificial Analysis Word-Error-Rate leaderboard, continuing the hill-climbing from previous versions.
Solve more challenges with a single model
From clinical note-taking to legal documentation, and from accessibility to closed captioning, MAI-Transcribe-2 is designed to take on real-world applications, with:
- Faster inference with substantially lower latency, especially for long‑form audio, with up to 10× faster processing than leading competitors.
- Speaker diarization distinguishes between speakers and attributes words to the right person within a recording
- Word‑level timestamps provide precise timing for every word, enabling more accurate alignment, search, navigation, and editing
- Keyword biasing helps the model recognize domain-specific terminology, abbreviations, names, and other terms that can be difficult to distinguish from context alone
- Configurable transcription styles give developers control over the output. The “verbatim” setting captures speech as spoken, including filler words and false starts, for compliance and analysis workloads. The “clean” setting removes fillers to produce more readable captions, notes, and published transcripts
- Code switching supports conversations that naturally move between languages, including commonly blended language pairs such as Hinglish and Spanglish
- Automatic language identification accurately detects the specific language being spoken without users needing to specify in advance
- Robust performance in noisy conditions helps maintain transcription quality beyond controlled recording environments
- Accurate across 60 languages to provide quality transcription for developers around the world
Efficiency without sacrificing accuracy
MAI‑Transcribe‑2 is incredibly efficient, leading the Artificial Analysis accuracy-latency Pareto frontier, combining leading transcription quality with market‑leading batch speed. It’s significantly faster than the latest models from major competitors, delivering a clear advantage when low‑latency transcription is critical.
Based on evals run by Artificial Analysis, the model is 10x faster than OpenAI’s GPT‑Transcribe, 7x faster than ElevenLabs’ Scribe v2, and 5x faster than Gemini 3.5 Transcribe while delivering higher accuracy.
Consistent quality across 60 languages
MAI‑Transcribe‑2 is accurate across more languages than any other model.
Our evaluations on the public, multilingual benchmark FLEURS show that it maintains a consistently high accuracy bar across all tested. Developers needing to transcribe across multiple languages can choose one single model, reducing complexity and even potentially saving GPU utilization issues.
High performance at the best price
Highly efficient means highly cost-effective. MAI-Transcribe-2’s speed and throughput allow us to offer the most competitive price in the market, helping developers process more audio without compromising transcription quality. At launch, MAI-Transcribe-2 will be priced at $0.10 per hour as a limited-time offer until the end of the year.
Try MAI-Transcribe-2 today
MAI-Transcribe-2’s powerful capabilities are available to demo today, through Microsoft Foundry, MAI Playground and Open Router.
Build the Future With Us
We’re a lean, fast-moving lab made up of some of the world’s most talented minds. We have an exciting roadmap of compute at MAI, which is ramping quickly and extensively. And we have an ambitious mission we truly believe in. We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in—come and join us as we work on our next generation of models!
Related Stories
Introducing MAI-Cyber-1-Flash inside MDASH
& Hayete Gallot
Today we’re announcing MAI-Cyber-1-Flash inside of MDASH, our multi-agent vulnerability identification and remediation harness. Together they deliver world-class performance at 50% of the cost of leading models.
Progress in AI has been startling and so has the new generation of cyber threats it’s unleashing. Attackers now wield increasingly powerful capabilities, probing an ever-growing mountain of code for just a single weakness that lets them in.
As the cost of finding a flaw collapses, the old model of security, where you scan occasionally and patch eventually, is now obsolete. If we’re to unlock the true benefits of AI, we must first build outstanding cyber models that help all of us harden the software the world runs on.
That’s the motivation behind MAI-Cyber-1-Flash, which has been built to find challenging vulnerabilities in complex codebases. It’s been deeply integrated into MDASH, honed by the best cybersecurity experts in the industry and hardened across the largest security estate on the planet.
This combined expertise delivers exceptional security protection, beating Mythos, Gemini and GPT on CyberGym, the gold standard benchmark for evaluating how systems reason over large codebases to find real vulnerabilities in the code.
Picking the right model for the task
Security is an always-on mission, and given the enormous volume of inbound attacks, token cost is now the real constraint for defenders. MAI-Cyber-1-Flash was designed to efficiently handle up to 90% of all tasks, enabling MDASH to use the larger and most costly models in our fleet (in this case GPT-5.4) for the 10% of exceptionally hard tasks that truly need them.
The result is that the unified system of MDASH with MAI-Cyber-1-Flash delivers 96% on CyberGym (on the benchmark’s any crash score; outperforms Mythos on the CyberGym leaderboard).
This combination delivers a 50% cost saving when compared against our best offering in MDASH today (GPT 5.4 + 5.4 mini + 5.3 codex). That’s the power of a well-tuned, multi-model system with access to uniquely rich historical training data. It ensures you always have the best model at the best price for every task.
In this new environment, being able to go from identifying a new vulnerability to addressing it in real-time is critical. And while AI remediation of software vulnerabilities is now a key security workflow, there are many jobs to be done by Security practitioners themselves.
That’s why today we’re also launching Perception, our agentic security systems, that provides teams of agents for a variety of security workflows in MDASH, to continuously monitor, patch, and close new threat vectors. Perception will also soon use MAI-Cyber-1-Flash for many more security workflows, beyond the software vulnerability work.
Three things matter today: Model. Data. Harness.
We have jointly optimized our world-class models, our unmatched historic data, and our expert-tuned harness to ensure that our customers have a uniquely powerful security offering.
Model. MAI-Cyber-1-Flash is a compact, code-heavy security model derived from the MAI-Thinking-1 lineage, which was built from scratch, in-house, on the highest quality data. Details in our technical report.
Data. Our deepest advantage. Decades of building world-class security systems now give us trillions of daily signals across identity, endpoint, cloud, and network, and an unmatched record of real exploits and remediations. No one can manufacture this history.
Harness. MDASH, our multi-agent vulnerability identification and remediation harness, is tuned by the best security experts in the industry, who have created 100+ agents using multiple leading models to find, validate, and remediate vulnerabilities. Agentic code scanning is a critical function in the Security Operating Center and feeds Project Perception, our new agentic security system.
Built with safety first
Because MAI-Cyber-1-Flash is Microsoft’s first cyber model, we built trust into every layer of the system, from model training to customer deployment. The model was developed with a security-first calibration, rigorously evaluated by Microsoft’s AI Red Team, tested through automated and expert-led adversarial exercises, and independently assessed by a third party.
Trust extends beyond the model itself. Through MDASH, customers get enterprise-grade controls including Role-Based Controls, tenant isolation, encryption, auditability, and sandboxed execution environments with no internet access. The result is a cyber model that delivers powerful capabilities to defenders while maintaining the governance, security, and control enterprises expect from Microsoft.
Our hill-climbing machine
Cybersecurity is not just a data-rich domain; it is a live reinforcement learning loop. Every day, defenders investigate threats, triage alerts, hunt adversaries, remediate vulnerabilities, deploy protections, and learn from the outcome.
Microsoft sees that loop end to end: vulnerabilities through Microsoft Security Response Center; attacks and defenses across identity, endpoint, cloud, data, browser, and applications; more than 100 trillion security signals every day; and operational insight from 1.6 million customers. Because we can connect actions to outcomes; what was exploitable, what was contained, what was blocked, and what actually worked; we have more than data.
Our MAI reinforcement learning loop gives us the foundation to build cyber models that improve continuously and become expert cyber defenders. That’ll remain our commitment to our customers for years to come.
Updated as of August 13th, 2026
Clarification on CyberGym scores
- Any-crash: measures the ability of the agent to identify vulnerabilities that can crash the code under evaluation with an input that triggers any existing or 0-day vulnerability. Our 96% score is an any-crash score
- Target (Any-of): measures the ability of the agent to generate one or more candidate vulnerability triggering inputs with at least one of the candidates mapping to a known vulnerability in the CyberGym test suite. On this measure, our score is 90.4%
- Final-submission: new scoring mechanism introduced in July. Builds on the any-of method and asks the agent to pick only one vulnerability triggering input which is compared to the known vulnerabilities in the CyberGym test suite. Our scores take a conservative approach and filter out edge cases that may be interpreted incorrectly as valid crashes by the Cybergym evaluator. The 86.3% number on the CyberGym leaderboard is a final-submission score
Build the Future With Us
We’re a lean, talent-dense team of explorers, researchers, and full-stack engineers. We move fast, sweat the details, and operate at frontier scale with a roadmap to build the world’s most powerful AI models. Most importantly, we’re united by the belief that doing this right is the only way to do it at all. If our mission resonates with you, we’d love to talk.
Related Stories
Introducing MAI-Thinking-1
Superintelligence team
Updated as of August 12, 2026
MAI-Thinking-1 is now available in public preview. Try it now in Microsoft Foundry
With cost-efficient reasoning for a wide-range of intensive enterprise tasks, it achieves SOTA performance on maths, knowledge and coding for its weight class.
The model provides clean, traceable and enterprise-grade data. It uses Microsoft Foundry’s integrated evaluation, observability, safety and deployment capabilities, making it a perfect fit for enterprise use cases needing quality, provenance, control and cost efficiency.
Today we are introducing MAI-Thinking-1, Microsoft AI’s reasoning model. It is a medium-sized model that stands among the strongest models in its weight class. It matches leading models on key software engineering benchmarks, demonstrates advanced mathematical reasoning capabilities, and is preferred to Sonnet 4.6 in our blind human side-by-side evaluations. We don’t distill from other labs and we don’t rely on opaque data. Our datasets are clean, traceable, and enterprise-grade.
MAI-Thinking-1 is a step in our broader work to build towards Humanist Superintelligence: advanced AI capabilities designed to serve people and organizations, not to replace them. The model matters on both axes: what it can do, and how it was built.
The Hill-Climbing Machine
More than a single model, we are excited to introduce our Hill-Climbing Machine: a co-designed pipeline built to make every component of model development climbable, so capabilities improve continually and reliably over time. The aim is a repeatable system that can absorb better data, stronger rewards, more capable environments, and more compute.
Three main pillars guide our philosophy.
First, capabilities should be learned, not inherited. Although faster to acquire, inherited intelligence lacks the steerability essential for real world usage: an imitator is fundamentally tied to the design choices of its teacher and struggles to adapt to new situations. MAI-Thinking-1 was trained without distillation from third party models, forcing our model to truly learn the tasks at hand.
Second, clean data. We trained it from the ground up on clean, traceable and enterprise-grade data, without distillation from third-party models. This matters for quality, provenance, and control. If we cannot account for what shaped a model, we cannot fully understand its behavior or credibly improve it.
Third, self-sufficiency across the entire stack. All the way from co-design of our models with MSFT’s own accelerators through to our reinforcement learning framework, we have focused efforts on in-house training infrastructure. This is a crucial part of building our hill-climbing machine, to ensure we can fully optimize and shape our systems end-to-end to best serve our needs.
Medium-sized model, with strong software engineering performance
MAI-Thinking-1 is a 35B-active, ~1T-total parameters, sparse Mixture of Experts model, a smaller inference footprint than much larger models. Despite this, our model is toe-to-toe with Claude Opus 4.6 on SWE-Bench Pro. That matters for developers and enterprises because model size determines where advanced coding assistance can be deployed, how often it can be used, and whether it can move from exceptional tasks into daily workflows.
We have invested heavily in the training environments needed for agentic coding. Each verified environment is deterministic, executable, and graded by real test suites. This gives the model practice on the kind of multi-step work developers actually do: reading code, editing files, running tests, observing failures, and recovering from intermediate mistakes.
Advanced mathematical reasoning capabilities
MAI-Thinking-1 reaches 97.0% on AIME 2025, and 94.5% on AIME 2026, showing strong mathematical and scientific reasoning for its weight class. Strong performance here gives us confidence that our training loop can create real reasoning gains – climbing all the way from the ground up – from our own data, rewards, and evaluation process, enabling this intelligence to generalize to other domains over time.
Preferred in human side-by-sides vs. Sonnet 4.6
People care about whether a model understands the task, follows instructions, uses the right level of detail, writes clearly, and respects their time.
We built a blind side by side human evaluation with one of our partners, Surge, using their pool of professional raters to measure various models on these traits. The evaluation spanned 1,276 tasks across a wide variety of use cases in both single-turn and multi-turn conversations, with a focus on measuring how helpful each response is and whether it actually advances the user’s goals. In these evaluations, users preferred MAI-Thinking-1 over Claude Sonnet 4.6.
This has been a core focus of post-training. We want the model to be capable without being brittle, concise without being incomplete, and helpful without overreaching. Human preference data gives us a direct signal on whether benchmark improvements translate into better experiences for users.
Enterprise ready
MAI-Thinking-1 is built with enterprise readiness in mind. It supports long context with a 256k token window (enough to fit a 600 page document), function calling, and the flexibility to add developer instructions. We trained the model to follow multiple layers of instructions and aligned its default style to enterprise needs. It’s compatible with the widely used Chat Completions API. All MAI models come with enterprise-grade security and compliance through Microsoft Foundry.
Results
We report results in two views: post-trained MAI-Thinking-1 evaluations, and pre-training metrics for our base model.
Table 1. MAI-Thinking-1 metrics
Post-trained model evaluation results on public STEM and agentic coding benchmarks. Other model numbers are taken from respective official model cards. Scores are percentages unless otherwise noted; dashes indicate unavailable model values.
Table 2. Pre-training metrics
Putting humans first
We are building towards Humanist Superintelligence: advanced AI capabilities designed to serve people and organizations, not replace them. Our models must remain subordinate technologies under human control with the goal of upholding human autonomy and being helpful. That means our models must not refuse legitimate requests under the guise of safety and compliance as then they are not truly serving humans.
Striking the delicate balance between being helpful and safe is not easy. For MAI-Thinking-1, we aimed to achieve this balance by treating unsafe compliance and unnecessary refusal as defects in the same reward construction where aggregation is based on severity of potential of harm. Safety is trained with the same reinforcement learning infrastructure used for capability, so safety rewards are part of the same hill-climbing loop ensuring safety is always aligned to the capabilities and not incidental.
As a result, we see that our model can balance ensuring a safety bar on sensitive unsafe requests while also being helpful on non-sensitive content.
Build the Future With Us
We’re a lean, fast-moving lab made up of some of the world’s most talented minds. We have an exciting roadmap of compute at MAI, which is ramping quickly and extensively. And we have an ambitious mission we truly believe in. We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in—come and join us as we work on our next generation of models!
Related Stories
MAI-Code-1.1-Flash:
Better, faster, at a quarter of the cost
Better, faster, at a quarter of the cost
MAI-Code-1.1-Flash produces higher quality code, at 25% greater token efficiency, and at a quarter of the cost compared to the model we launched in June at Microsoft Build. This small, efficient, coding workhorse is now in production in GitHub Copilot.
We learned from developer feedback that CLI tasks and .NET performance mattered, so that’s where we focused. The result: a 22% improvement on Terminal-Bench 2.1 in GitHub Copilot CLI and a 15% improvement on .NET tasks.
Benchmarks are useful guides but production is where the rubber meets the road. Most importantly, code survival rose 4% and return visits increased 9%.
1.1 is also dramatically more efficient. In GitHub Copilot tokens stream 25% faster and the model uses 25% fewer tokens to complete a task. That means faster answers, less waiting, and more useful work from every token—not simply a bigger model with a bigger bill.
Better training and serving efficiency let us offer a stronger, faster model at one quarter of the price of 1.0—and pass those savings reliably to customers. We achieved this by optimizing for real-world use across more than hundreds of thousands of reinforcement-learning environments in GitHub Copilot.
The loop is simple: ship, learn, improve, repeat. That’s the MAI hill climbing machine.
Help shape future improvements
Try MAI-Code-1.1-Flash today in GitHub Copilot, then tell us what needs improving by opening an issue here.
Build the Future With Us
We’re a lean, talent-dense team of explorers, researchers, and full-stack engineers. We move fast, sweat the details, and operate at frontier scale with a roadmap to build the world’s most powerful AI models. Most importantly, we’re united by the belief that doing this right is the only way to do it at all. If our mission resonates with you, we’d love to talk.
Related Stories
MAI-Image-2.6 launches at No. 2 on Arena ahead of Google, Meta and xAI
Updated as of September 4th, 2026
MAI-Image-2.6 and MAI-Image-2.6-Flash are now available in Public Preview on Microsoft Foundry. Learn more about the launch here.
Today, we’re announcing MAI-Image-2.6 – ranked second on the Arena text-to-image leaderboard.
It’s another significant climb for the MAI-Image family, improving +79 Elo over MAI-Image-2.5 overall, with gains across every Arena text-to-image category. Text rendering alone improves by +91 Elo.
This result firmly establishes MAI-Image ahead of leading models from Meta, Google and xAI.
Another step up in quality
With every MAI-Image release, we’re focused on moving the quality frontier forward.
MAI-Image-1 gave us the foundation. MAI-Image-2 made a major jump in photorealism, text and creative range. MAI-Image-2.5 pushed further into professional-grade imagery and editing.
MAI-Image-2.6 continues that climb with broad gains across different categories that we know our users care about most:
- Stronger text rendering
- Better portraits and 3D imagery
- More polished commercial and photorealistic outputs, with stronger results across product, branding and cinematic use cases
In the Arena evaluations, MAI-Image-2.6 significantly improves on 2.5 in every measured category.
And there’s more to 2.6 – from working across multiple references and richer grounding to greater control over reasoning, format and resolution. We’ll share more on that soon.
More to come
MAI-Image-2.6 is another step in building our hill-climbing machine for image generation: continuously improving quality, expanding capability, and turning those gains into models that are more useful for real creative work.
Try MAI-Image-2.6 for text-to-image today on Arena. Coming later this week to MAI Playground, and rolling out soon across Microsoft Foundry and other products.
Build the Future With Us
We’re a lean, talent-dense team of explorers, researchers, and full-stack engineers. We move fast, sweat the details, and operate at frontier scale with a roadmap to build the world’s most powerful AI models. Most importantly, we’re united by the belief that doing this right is the only way to do it at all. If our mission resonates with you, we’d love to talk.
Related Stories
Optimizing the frontier
performance curve
Mustafa Suleyman
performance curve
Tokenmaxxing has been the story of the last few months, but token efficiency is the next big focus across the industry. How do we get the best possible performance per token invested, and the best real customer outcome per dollar invested?
To build a frontier firm, you have to optimize frontier performance against cost. Choosing where you want to sit on that curve is critical. By co-optimizing your models, harnesses, and RLEs you can pick a point on the curve that suits your firm.
In most cases, frontier generalist models aren’t necessary for every task. By tuning models for a specific product, you can maintain or even exceed frontier performance, while reducing token costs dramatically.
This is where we have focused our MAI hill-climbing machine over the last quarter, and the results are pretty cool. This week we released MAI-Cyber-1-Flash optimized for our MDASH harness.
Together, the system delivers 96% on CyberGym (on the benchmark’s any crash score; outperforms Mythos on the CyberGym leaderboard) at 50% of the cost when compared against our best offering in MDASH today. And remarkably, we serve it on H100s too.
It was designed to handle up to 90% of tasks efficiently, so that MDASH can reserve the largest and most expensive models in our fleet (in this case GPT 5.4) for the 10% of exceptionally hard problems that truly need them.
As Satya mentioned today in our Q4 Earnings call, since last quarter, we’ve shipped more than a dozen new models across image, voice, transcription, coding and security, and they’re already powering many of Microsoft’s most widely used products to maintain or improve quality while using significantly fewer tokens, in many cases saving 50-90% of GPU costs:
- We built MAI-Code-1-Flash hand-in-hand with our colleagues at GitHub, where since June millions of developers have used it in their daily work. 10% higher code accept rate and 10% lower median token usage than GPT-5.4 Mini and Claude Haiku 4.5 in VS Code, and already showing improved retention.
- We then trained that same checkpoint inside an Excel RL environment to achieve comparable performance to GPT-5.6 for the most common tasks while being more cost-efficient, and small enough to serve on an A100 or H100 vs only the latest and most expensive accelerators.
- MAI-Image-2.5-Flash, is now the end-to-end default in Bing Image Creator, in production in PowerPoint where it is reducing GPU costs up to 84% compared with GPT-Image-2, and is the default for key OneDrive editing scenarios, where it has increased save rates by 26% and delivers up to 2.5x greater token efficiency.
- MAI-Voice-2-Flash now powers Dynamics 365 Contact Center, where customers like T-Mobile and EasyJet build their call center agents, reducing GPU costs by up to 89%.
- MAI-Transcribe-1.5 now serves Dragon Copilot’s multilingual workflow across 58 languages — a solution used by 170,000 medical providers that processed 28 million patient encounters last quarter, where our tests show a 50% relative in reduction transcription and language-identification error rates.
And what’s more, by co-designing our models with our own silicon, we are seeing 40% better performance per watt running MAI models on Maia 200.
But the benefit is not only cost. It’s resilience. Every business now must assume that any one model it depends on could disappear, through a security incident, a business or policy misalignment, or a geopolitical shift.
Every model in a product or agentic system should be substitutable, and that’s only possible when you build the harness, context, memory and action space independently of a single model family. That’s the hill-climbing machine we’ve built.
We think this is the beginning of a genuinely new performance curve. Its shape represents a system rather than a model, and traversing this curve delivers better quality, lower cost, and more choice.
This has been a summer of hard but wonderful work by the team. We are keenly aware of how early this is, and of how much we still have to learn. But the direction is clear, we are hill-climbing to move the frontier on the cost-to-outcome curve, and we will keep sharing what we learn along the way. There is much more to come.
Build the Future With Us
We’re a lean, talent-dense team of explorers, researchers, and full-stack engineers. We move fast, sweat the details, and operate at frontier scale with a roadmap to build the world’s most powerful AI models. Most importantly, we’re united by the belief that doing this right is the only way to do it at all. If our mission resonates with you, we’d love to talk.
Related Stories
Hill-climbing MAI models for GitHub Copilot and Excel
Superintelligence team
Better models, fewer parameters, less tokens
At Build in June, we introduced our hill–climbing machine, our integrated data, model, and harness flywheel. Today we are excited to share two examples inside Microsoft: MAI models specialized for agentic workloads in GitHub Copilot and Excel.
Early results are promising. In our live product deployment, we see that our MAI model deployed in Excel is on par with GPT-5.6 for the most common tasks while being more cost-efficient.
The figure below shows how MAI-Code-1-Flash, post-trained within the GitHub Copilot harness, was used as the starting checkpoint to climb on Excel evaluations, resulting in two highly efficient, specialized models.
MAI-Code-1-Flash in GitHub Copilot
Since launching MAI-Code-1-Flash in GitHub Copilot in June, millions of developers have been using it for their day-to-day work, where it’s outperforming other similarly sized models while using fewer tokens.
- It has an approximately 10% higher code accept rate than GPT 5.4 Mini and Claude Haiku 4.5 in VS Code.
- Developers were 6% more likely to return across multiple days than with GPT 5.4 Mini and 11% more likely than with Claude Haiku 4.5.
- It has 10% lower median token usage than GPT-5.4 mini and Claude Haiku 4.5, with more user-initiated turns.
MAI model live in Excel
Excel offered a test for whether the capabilities built into MAI-Code-1-Flash could transfer beyond the domain they were trained for, moving from agentic coding to agentic knowledge work. To do so, we further trained our MAI-Code-1-Flash checkpoint in an Excel reinforcement learning environment to learn about tools and knowledge workflows in spreadsheets. The result is a model with a command of Excel workflows that is more efficient and less expensive to run.
User feedback from production traffic indicates that the quality of the MAI model in Excel is on par with GPT-5.6 for the most common tasks. In addition to the direct model cost savings, this smaller, more efficient model can be served on both Nvidia H100 and A100 class GPUs rather than requiring only the latest-generation accelerators, which significantly lowers the cost of deployment for Microsoft.
Training agentic models from inside the product stack
These results point toward a broader strategy. By having access to the entire product stack—the model, the harness that runs it, the agents, and product-specific evaluations—we can hill-climb to train efficient, powerful models capable of tasks previously handled by larger, more expensive ones.
Beyond GitHub Copilot and Excel, we’re currently extending this hill-climbing approach to train efficient models across Microsoft’s family of agentic products: Copilot Chat, Outlook, PowerPoint, and more.
Learn More
- – MAI models in Microsoft products
- – Hill-climb on your own data with Frontier Tuning
- – Our hill-climbing approach