Introducing MAI-Voice-2

June 2, 2026
Models
Superintelligence team

Today we’re launching MAI-Voice-2 — the most expressive, natural-sounding text-to-speech model we’ve built to date. It’s a significant leap from its predecessor across every dimension that matters to production voice experiences: fidelity, language coverage, speaker consistency, and emotional range. It is built for the products and services where voice quality directly impacts user experience: assistants or customer support that represent your brand, audiobooks that hold attention over hours, and accessibility experiences where voice is the only interface. It’s also built with responsible deployment in mind, with consent guardrails ensuring the technology is as trustworthy as it sounds. MAI-Voice-2 is now available in Microsoft Foundry, and is being integrated into VSCode and the Dynamics 365 Contact Center.

Features and capabilities

  • Expanding from English‑only to 15 languages while maintaining the same naturalness and expressiveness as English.
  • Granular emotion control via emotion tags: sad, whispered, excited, etc.
  • Zero-shot voice prompting using 5-60s of reference audio available for all supported languages, with built-in consent guardrails.
  • MAI-Voice-2 is preferred over its predecessor MAI-Voice-1 72% of the time.
  • Stable speaker identity across long-form content – audiobooks, podcasts, lectures.
  • Code-switching capabilities for select language pairs — such as Hindi-English and Spanish-English — matching the way users naturally mix languages in everyday speech.

Hear it for yourself:

English (emotion: Embarrassed)

So I was just standing there, right? And then (sigh) oh my God, she actually said it to his face. I mean, honestly, good for her.

German (emotion: Confused)

Häh? Warum schicken dir mir eine Mahnung? Das macht keinen Sinn. Ich hab das doch schon vor zwei Wochen bezahlt.

Hindi (emotion: Excited)

अरे यार धीरे बोल, कोई सुन लेगा तो पूरा surprise ही लीक हो जाएगा! इतने साल बाद मुंबई में उससे मिलने वाला हूँ.दिल full Bollywood-mode में है

English (role: Motivational Trainer)

Alright, time to focus. Notice how the egret doesn’t rush the moment, it studies it. Every movement is deliberate, every pause intentional. That’s discipline. That’s control. So when the opportunity appears, you can strike without hesitation. Patience earns the catch.

English (role: Sports Commentator)

With everything on the line, the egret makes its move! Slow through the shallows… watching… waiting… And it’s a sudden strike! Got it! Incredible precision from the long beak! The fish never saw it coming. What a scene! Complete composure under pressure. A masterclass performance here in the pond tonight.

Performance

MAI-Voice-2 generates very natural speech in a controllable way. In side-by-side preference tests, it was preferred over its predecessors 72% of the time. In speaker similarity evaluations, speech generated by MAI-Voice-2 is indistinguishable from recordings of the same voice. Below, you can verify this yourself by trying to identify where the human speech ends and the MAI-Voice-2 output begins.

Bar chart showing MAI-Voice-2 with a 72.1% win rate and MAI-Voice-1 with a 27.9% win rate for overall quality preference out of 2,500 listening tests.
Bar graph showing that, on average across 11 languages, 45.5% of listeners preferred MAI-Voice-2 generated speech, 44% preferred real human recordings, and 10.5% resulted in a tie, out of 2,222 responses.

Guess the human recording vs. MAI‑Voice‑2

Listen to the audio clips below – each blends human recordings with speech generated by MAI‑Voice‑2. Can you tell where the human voice ends and the synthetic voice begins, or vice versa? Or does it sound like one continuous voice?

Human recorded + TTS

Language: English US

Human recorded + TTS

Language: Hindi (India)

Human recorded + TTS

Language: Spanish (Mexico)

TTS + Human recorded

Language: French (France)

Human recorded + TTS

Language: German (Germany)

Supported Languages

We prioritized depth across 15 languages, ensuring for supported languages we support a spectrum of expressive capabilities spanning tonal, pitch accent, stress timed, and syllable timed systems. We plan to continue expanding and refining the expressive range for all supported languages.

MAI-Voice-2 now supports the following languages/locales: English (US), English (Australia), Italian, French, German, Hindi, Spanish (Spain), Spanish (Mexico), Portuguese (Brazil), Portuguese (Portugal), Korean, Chinese (Simplified), Turkish, Russian, Thai, Dutch, Romanian and Hungarian.

In markets where people naturally mix languages, we support code-switching – notably Hindi–English and Spanish–English – reflecting how people actually speak. In internal testing, the model switches languages mid sentence fluidly, without losing prosodic naturalness nor speaker identity.

Hindi + English

Oh my god, just look at this gorgeous sunset! क्या तुमने कभी ऐसा beautiful sky देखा है? It looks just like a painting, with all these stunning colours… गुलाबी, नारंगी, बैंगनी। It’s literally magical


Spanish (Mexican) + English

Quesadillas, tacos, enchiladas, y guacamole are staples of Mexican cuisine, pero también incluyen ingredients like cilantro, jalapeños, and queso fresco for authentic, traditional, regional preparations.


Voice Synthesis

Developers can create a custom voice in Microsoft Foundry across all supported languages using just a short reference clip – no retraining or fine tuning required. With only a few seconds of audio (recommended: 5–60 seconds), MAI Voice 2 can generate high quality speech that matches the speaker’s identity, making it easy for companies to bring their own brand voice into products without maintaining a separate voice model.

Consent and Safety

Consent is enforced at the system level: only authorized, licensed voices can be synthesized in production. No unlicensed voice cloning is possible. To gain access to this feature apply here.

Use Cases

  • Assistants: Branded voices for Copilot, apps, devices, customer support.
  • Entertainment: Characters for games, podcasts, audiobooks, AR/VR.
  • Accessibility: Narration for visually impaired users; voice for speech impairments.
  • Education: Instructors and characters for courses and simulations.
  • Creators: Turn text into audio with your own voice. No studio required.

Try it out

DuoAI

DuoAI is an experimental experience that gives you a direct way to try MAI‑Voice‑2, MAI‑Transcribe‑1.5, and MAI‑Image‑2.5 models in action – showcasing natural, fluid, expressive dialogue. In the demo, you can engage in a three‑way conversation with two agents and even generate images using MAI‑Image‑2.5. It’s a practical preview of how MAI multimodal models work together to build powerful, customizable voice agents. Try DuoAI now

Note: DuoAI is not meant to showcase the capabilities of the underlying LLM – that component is modular and can be swapped as needed.

You can also explore the models directly in the MAI Playground.

Learn more about MAI-Voice-2

  • Model card [Link]
  • Foundry API documentation [Link]
  • Cookbook [Link]

Build the Future With Us

We’re a lean, fast-moving lab made up of some of the world’s most talented minds. We have an exciting roadmap of compute at MAI, with our next-generation GB200 cluster now operational. And we have an ambitious mission we truly believe in. We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in—come and join us as we work on our next generation of models!

Explore all jobs

Related Stories

Pushing the quality-cost frontier with MAI-Image-2.6

September 4, 2026
Models
Three painted flowers—pink, purple, and yellow—on thin, curved stems with green leaves, set against a plain blue background.

MAI-Image-2.6 is our strongest image model yet. Today, we’re bringing it to developers in Microsoft Foundry and expanding the family with MAI-Image-2.6-Flash – built to deliver that same level of quality for latency-sensitive, high-throughput production workloads.

Both models come with support for multi-image reference editing, web grounding, and dynamic aspect ratios. Developers now have choice between maximum precision with MAI-Image-2.6 and production speed with MAI-Image-2.6-Flash.

Quality at production speed

MAI-Image-2.6 has established itself among the industry’s leading image models. Today, it ranks No. 2 for both text to image and image editing on Arena.[1] On Artificial Analysis, it ranks No. 2 for text-to-image and No. 1 for image editing.[2]

MAI-Image-2.6-Flash brings comparable quality to latency-sensitive, high-throughput workloads. It is able to generate images 2.8x faster than GPT-Image-2-Medium while delivering 72% greater efficiency.

Leading quality for the price

MAI-Image-2.6’s combination of quality and efficiency delivers the best price-per-Elo performance in the world, helping production teams scale high-quality image generation while keeping token usage and costs under control.

More ways to create

We’ve been listening closely to how users create with our image models and have expanded their capabilities around the creative workflows that matter most.

The latest models offer more control, more context, and more ways to bring ideas to life:

  • Multi-reference editing brings together people, products, styles, and scenes from different images
  • Web grounding pulls info from across the web to create rich visuals informed by relevant, up-to-date information
  • Higher resolutions and dynamic aspect ratios chooses the optimal format for compositions with support for up to 1.5K resolution.

Start Creating

Try both models in MAI Playground, or start building in Public Preview through Microsoft Foundry.

FOOTNOTE:
[1] As of Sep 4, 2026
[2] As of Sep 4, 2026

Build the Future With Us

We’re a lean, fast-moving lab made up of some of the world’s most talented minds. We have an exciting roadmap of compute at MAI, which is ramping quickly and extensively. And we have an ambitious mission we truly believe in. We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in—come and join us as we work on our next generation of models!

Explore all jobs

Related Stories

MAI-Transcribe-2 is the fastest, most accurate and cheapest speech recognition model in the world

September 3, 2026
Models
Abstract image with horizontal green, yellow, and pink bands, evoking the feel of a blurry landscape or seascape at sunset, with a soft, hazy, and gradient effect throughout.

Introducing MAI‑Transcribe‑2. It’s not only our most capable transcription model yet, but the most capable and efficient amongst our competitors.

With new features like diarization, configurable transcription styles, and word-level timestamps, MAI-Transcribe-2 beats other leading models like Gemini 3.5 Transcribe, GPT-Transcribe, Whisper V3-Large, and ScribeV2, while also handling a broader range of real‑world audio.

Our model ranks first on the FLEURS benchmark across 60 languages with an average Word-Error-Rate of 5.2%, defines the Pareto Frontier for accuracy and latency on Artificial Analysis, and ranks second on the Artificial Analysis Word-Error-Rate leaderboard, continuing the hill-climbing from previous versions.

All this performance also comes at the best price on the market, just $0.10 per hour of audio.

Solve more challenges with a single model

From clinical note-taking to legal documentation, and from accessibility to closed captioning, MAI-Transcribe-2 is designed to take on real-world applications, with:

  • Faster inference with substantially lower latency, especially for long‑form audio, with up to 10× faster processing than leading competitors.
  • Speaker diarization distinguishes between speakers and attributes words to the right person within a recording
  • Word‑level timestamps provide precise timing for every word, enabling more accurate alignment, search, navigation, and editing
  • Keyword biasing helps the model recognize domain-specific terminology, abbreviations, names, and other terms that can be difficult to distinguish from context alone
  • Configurable transcription styles give developers control over the output. The “verbatim” setting captures speech as spoken, including filler words and false starts, for compliance and analysis workloads. The “clean” setting removes fillers to produce more readable captions, notes, and published transcripts
  • Code switching supports conversations that naturally move between languages, including commonly blended language pairs such as Hinglish and Spanglish
  • Automatic language identification accurately detects the specific language being spoken without users needing to specify in advance
  • Robust performance in noisy conditions helps maintain transcription quality beyond controlled recording environments
  • Accurate across 60 languages to provide quality transcription for developers around the world

Efficiency without sacrificing accuracy

MAI‑Transcribe‑2 is incredibly efficient, leading the Artificial Analysis accuracy-latency Pareto frontier, combining leading transcription quality with market‑leading batch speed. It’s significantly faster than the latest models from major competitors, delivering a clear advantage when low‑latency transcription is critical.

Based on evals run by Artificial Analysis, the model is 10x faster than OpenAI’s GPT‑Transcribe, 7x faster than ElevenLabs’ Scribe v2, and 5x faster than Gemini 3.5 Transcribe while delivering higher accuracy.

Consistent quality across 60 languages

MAI‑Transcribe‑2 is accurate across more languages than any other model.

Our evaluations on the public, multilingual benchmark FLEURS show that it maintains a consistently high accuracy bar across all tested. Developers needing to transcribe across multiple languages can choose one single model, reducing complexity and even potentially saving GPU utilization issues.

High performance at the best price

Highly efficient means highly cost-effective. MAI-Transcribe-2’s speed and throughput allow us to offer the most competitive price in the market, helping developers process more audio without compromising transcription quality. At launch, MAI-Transcribe-2 will be priced at $0.10 per hour as a limited-time offer until the end of the year.

Try MAI-Transcribe-2 today

MAI-Transcribe-2’s powerful capabilities are available to demo today, through Microsoft Foundry, MAI Playground and Open Router.



Build the Future With Us

We’re a lean, fast-moving lab made up of some of the world’s most talented minds. We have an exciting roadmap of compute at MAI, which is ramping quickly and extensively. And we have an ambitious mission we truly believe in. We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in—come and join us as we work on our next generation of models!

Explore all jobs

Related Stories

MAI-Image-2.6 launches at No. 2 on Arena ahead of Google, Meta and xAI

August 10, 2026
Models

Updated as of September 4th, 2026

MAI-Image-2.6 and MAI-Image-2.6-Flash are now available in Public Preview on Microsoft Foundry. Learn more about the launch here.

A large number 2 with a hashtag, filled with small photos of nature, animals, and people, overlaid on faded text about rendering, imaging, modeling, and art.

Today, we’re announcing MAI-Image-2.6 – ranked second on the Arena text-to-image leaderboard.

It’s another significant climb for the MAI-Image family, improving +79 Elo over MAI-Image-2.5 overall, with gains across every Arena text-to-image category. Text rendering alone improves by +91 Elo.

This result firmly establishes MAI-Image ahead of leading models from Meta, Google and xAI.

Another step up in quality

With every MAI-Image release, we’re focused on moving the quality frontier forward.

MAI-Image-1 gave us the foundation. MAI-Image-2 made a major jump in photorealism, text and creative range. MAI-Image-2.5 pushed further into professional-grade imagery and editing.

MAI-Image-2.6 continues that climb with broad gains across different categories that we know our users care about most:

  • Stronger text rendering
  • Better portraits and 3D imagery
  • More polished commercial and photorealistic outputs, with stronger results across product, branding and cinematic use cases

In the Arena evaluations, MAI-Image-2.6 significantly improves on 2.5 in every measured category.

And there’s more to 2.6 – from working across multiple references and richer grounding to greater control over reasoning, format and resolution. We’ll share more on that soon.

More to come

MAI-Image-2.6 is another step in building our hill-climbing machine for image generation: continuously improving quality, expanding capability, and turning those gains into models that are more useful for real creative work.

Try MAI-Image-2.6 for text-to-image today on Arena. Coming later this week to MAI Playground, and rolling out soon across Microsoft Foundry and other products.

Build the Future With Us

We’re a lean, talent-dense team of explorers, researchers, and full-stack engineers. We move fast, sweat the details, and operate at frontier scale with a roadmap to build the world’s most powerful AI models. Most importantly, we’re united by the belief that doing this right is the only way to do it at all. If our mission resonates with you, we’d love to talk.

Explore all jobs

Related Stories

MAI-Code-1.1-Flash:
Better, faster, at a quarter of the cost 

August 11, 2026
Models
Two large white curly brackets on a green background overlayed with rows of faint text and code-like characters, creating a digital, abstract effect.

MAI-Code-1.1-Flash produces higher quality code, at 25% greater token efficiency, and at a quarter of the cost compared to the model we launched in June at Microsoft Build. This small, efficient, coding workhorse is now in production in GitHub Copilot.

We learned from developer feedback that CLI tasks and .NET performance mattered, so that’s where we focused. The result: a 22% improvement on Terminal-Bench 2.1 in GitHub Copilot CLI and a 15% improvement on .NET tasks.

Benchmarks are useful guides but production is where the rubber meets the road. Most importantly, code survival rose 4% and return visits increased 9%.

1.1 is also dramatically more efficient. In GitHub Copilot tokens stream 25% faster and the model uses 25% fewer tokens to complete a task. That means faster answers, less waiting, and more useful work from every token—not simply a bigger model with a bigger bill.

Better training and serving efficiency let us offer a stronger, faster model at one quarter of the price of 1.0—and pass those savings reliably to customers. We achieved this by optimizing for real-world use across more than hundreds of thousands of reinforcement-learning environments in GitHub Copilot.

The loop is simple: ship, learn, improve, repeat. That’s the MAI hill climbing machine.

Help shape future improvements

Try MAI-Code-1.1-Flash today in GitHub Copilot, then tell us what needs improving by opening an issue here.

Build the Future With Us

We’re a lean, talent-dense team of explorers, researchers, and full-stack engineers. We move fast, sweat the details, and operate at frontier scale with a roadmap to build the world’s most powerful AI models. Most importantly, we’re united by the belief that doing this right is the only way to do it at all. If our mission resonates with you, we’d love to talk.

Explore all jobs

Related Stories

Optimizing the frontier
performance curve

July 29, 2026
Models
Mustafa Suleyman
Nine app icons, including mail, JetBrains, GitHub, PowerPoint, Excel, Visual Studio Code, and Microsoft OneDrive, displayed in a row on a blurred green and pink gradient background with dotted lines.

Tokenmaxxing has been the story of the last few months, but token efficiency is the next big focus across the industry. How do we get the best possible performance per token invested, and the best real customer outcome per dollar invested?

To build a frontier firm, you have to optimize frontier performance against cost. Choosing where you want to sit on that curve is critical. By co-optimizing your models, harnesses, and RLEs you can pick a point on the curve that suits your firm.

In most cases, frontier generalist models aren’t necessary for every task. By tuning models for a specific product, you can maintain or even exceed frontier performance, while reducing token costs dramatically.

This is where we have focused our MAI hill-climbing machine over the last quarter, and the results are pretty cool. This week we released MAI-Cyber-1-Flash optimized for our MDASH harness.

Together, the system delivers 96% on CyberGym (on the benchmark’s any crash score; outperforms Mythos on the CyberGym leaderboard) at 50% of the cost when compared against our best offering in MDASH today. And remarkably, we serve it on H100s too.

It was designed to handle up to 90% of tasks efficiently, so that MDASH can reserve the largest and most expensive models in our fleet (in this case GPT 5.4) for the 10% of exceptionally hard problems that truly need them.

As Satya mentioned today in our Q4 Earnings call, since last quarter, we’ve shipped more than a dozen new models across image, voice, transcription, coding and security, and they’re already powering many of Microsoft’s most widely used products to maintain or improve quality while using significantly fewer tokens, in many cases saving 50-90% of GPU costs:

  • We built MAI-Code-1-Flash hand-in-hand with our colleagues at GitHub, where since June millions of developers have used it in their daily work. 10% higher code accept rate and 10% lower median token usage than GPT-5.4 Mini and Claude Haiku 4.5 in VS Code, and already showing improved retention.
  • We then trained that same checkpoint inside an Excel RL environment to achieve comparable performance to GPT-5.6 for the most common tasks while being more cost-efficient, and small enough to serve on an A100 or H100 vs only the latest and most expensive accelerators.
  • MAI-Image-2.5-Flash, is now the end-to-end default in Bing Image Creator, in production in PowerPoint where it is reducing GPU costs up to 84% compared with GPT-Image-2, and is the default for key OneDrive editing scenarios, where it has increased save rates by 26% and delivers up to 2.5x greater token efficiency.
  • MAI-Voice-2-Flash now powers Dynamics 365 Contact Center, where customers like T-Mobile and EasyJet build their call center agents, reducing GPU costs by up to 89%.
  • MAI-Transcribe-1.5 now serves Dragon Copilot’s multilingual workflow across 58 languages — a solution used by 170,000 medical providers that processed 28 million patient encounters last quarter, where our tests show a 50% relative in reduction transcription and language-identification error rates.

And what’s more, by co-designing our models with our own silicon, we are seeing 40% better performance per watt running MAI models on Maia 200.

But the benefit is not only cost. It’s resilience. Every business now must assume that any one model it depends on could disappear, through a security incident, a business or policy misalignment, or a geopolitical shift.

Every model in a product or agentic system should be substitutable, and that’s only possible when you build the harness, context, memory and action space independently of a single model family. That’s the hill-climbing machine we’ve built.

We think this is the beginning of a genuinely new performance curve. Its shape represents a system rather than a model, and traversing this curve delivers better quality, lower cost, and more choice.

This has been a summer of hard but wonderful work by the team. We are keenly aware of how early this is, and of how much we still have to learn. But the direction is clear, we are hill-climbing to move the frontier on the cost-to-outcome curve, and we will keep sharing what we learn along the way. There is much more to come.

Build the Future With Us

We’re a lean, talent-dense team of explorers, researchers, and full-stack engineers. We move fast, sweat the details, and operate at frontier scale with a roadmap to build the world’s most powerful AI models. Most importantly, we’re united by the belief that doing this right is the only way to do it at all. If our mission resonates with you, we’d love to talk.

Explore all jobs

Related Stories

Introducing MAI-Cyber-1-Flash inside MDASH

World-class security at half the cost
August 13, 2026
Models
Mustafa Suleyman
& Hayete Gallot
Abstract illustration of three overlapping shield shapes in blue, pink, and purple tones on a beige background.

Today we’re announcing MAI-Cyber-1-Flash inside of MDASH, our multi-agent vulnerability identification and remediation harness. Together they deliver world-class performance at 50% of the cost of leading models.

Progress in AI has been startling and so has the new generation of cyber threats it’s unleashing. Attackers now wield increasingly powerful capabilities, probing an ever-growing mountain of code for just a single weakness that lets them in.

As the cost of finding a flaw collapses, the old model of security, where you scan occasionally and patch eventually, is now obsolete. If we’re to unlock the true benefits of AI, we must first build outstanding cyber models that help all of us harden the software the world runs on.

That’s the motivation behind MAI-Cyber-1-Flash, which has been built to find challenging vulnerabilities in complex codebases. It’s been deeply integrated into MDASH, honed by the best cybersecurity experts in the industry and hardened across the largest security estate on the planet.

This combined expertise delivers exceptional security protection, beating Mythos, Gemini and GPT on CyberGym, the gold standard benchmark for evaluating how systems reason over large codebases to find real vulnerabilities in the code.

Bar chart titled "CyberGym Evaluation" comparing success rates of five models, with MDASH: MAI-Cyber-1-Flash + GPT-5.4 leading at 95.95%, and other four models ranging between 83.2% and 85.6%.

Picking the right model for the task

Security is an always-on mission, and given the enormous volume of inbound attacks, token cost is now the real constraint for defenders. MAI-Cyber-1-Flash was designed to efficiently handle up to 90% of all tasks, enabling MDASH to use the larger and most costly models in our fleet (in this case GPT-5.4) for the 10% of exceptionally hard tasks that truly need them.

The result is that the unified system of MDASH with MAI-Cyber-1-Flash delivers 96% on CyberGym (on the benchmark’s any crash score; outperforms Mythos on the CyberGym leaderboard).

This combination delivers a 50% cost saving when compared against our best offering in MDASH today (GPT 5.4 + 5.4 mini + 5.3 codex). That’s the power of a well-tuned, multi-model system with access to uniquely rich historical training data. It ensures you always have the best model at the best price for every task.

In this new environment, being able to go from identifying a new vulnerability to addressing it in real-time is critical. And while AI remediation of software vulnerabilities is now a key security workflow, there are many jobs to be done by Security practitioners themselves.

That’s why today we’re also launching Perception, our agentic security systems, that provides teams of agents for a variety of security workflows in MDASH, to continuously monitor, patch, and close new threat vectors. Perception will also soon use MAI-Cyber-1-Flash for many more security workflows, beyond the software vulnerability work.

Three things matter today: Model. Data. Harness.

We have jointly optimized our world-class models, our unmatched historic data, and our expert-tuned harness to ensure that our customers have a uniquely powerful security offering.

Model. MAI-Cyber-1-Flash is a compact, code-heavy security model derived from the MAI-Thinking-1 lineage, which was built from scratch, in-house, on the highest quality data. Details in our technical report.

Data. Our deepest advantage. Decades of building world-class security systems now give us trillions of daily signals across identity, endpoint, cloud, and network, and an unmatched record of real exploits and remediations. No one can manufacture this history.

Harness. MDASH, our multi-agent vulnerability identification and remediation harness, is tuned by the best security experts in the industry, who have created 100+ agents using multiple leading models to find, validate, and remediate vulnerabilities. Agentic code scanning is a critical function in the Security Operating Center and feeds Project Perception, our new agentic security system.

Built with safety first

Because MAI-Cyber-1-Flash is Microsoft’s first cyber model, we built trust into every layer of the system, from model training to customer deployment. The model was developed with a security-first calibration, rigorously evaluated by Microsoft’s AI Red Team, tested through automated and expert-led adversarial exercises, and independently assessed by a third party.

Trust extends beyond the model itself. Through MDASH, customers get enterprise-grade controls including Role-Based Controls, tenant isolation, encryption, auditability, and sandboxed execution environments with no internet access. The result is a cyber model that delivers powerful capabilities to defenders while maintaining the governance, security, and control enterprises expect from Microsoft.

Our hill-climbing machine

Cybersecurity is not just a data-rich domain; it is a live reinforcement learning loop. Every day, defenders investigate threats, triage alerts, hunt adversaries, remediate vulnerabilities, deploy protections, and learn from the outcome.

Microsoft sees that loop end to end: vulnerabilities through Microsoft Security Response Center; attacks and defenses across identity, endpoint, cloud, data, browser, and applications; more than 100 trillion security signals every day; and operational insight from 1.6 million customers. Because we can connect actions to outcomes; what was exploitable, what was contained, what was blocked, and what actually worked; we have more than data.

Our MAI reinforcement learning loop gives us the foundation to build cyber models that improve continuously and become expert cyber defenders. That’ll remain our commitment to our customers for years to come.

Updated as of August 13th, 2026

Clarification on CyberGym scores

  • Any-crash: measures the ability of the agent to identify vulnerabilities that can crash the code under evaluation with an input that triggers any existing or 0-day vulnerability. Our 96% score is an any-crash score
  • Target (Any-of): measures the ability of the agent to generate one or more candidate vulnerability triggering inputs with at least one of the candidates mapping to a known vulnerability in the CyberGym test suite. On this measure, our score is 90.4%
  • Final-submission: new scoring mechanism introduced in July. Builds on the any-of method and asks the agent to pick only one vulnerability triggering input which is compared to the known vulnerabilities in the CyberGym test suite. Our scores take a conservative approach and filter out edge cases that may be interpreted incorrectly as valid crashes by the Cybergym evaluator. The 86.3% number on the CyberGym leaderboard is a final-submission score

Build the Future With Us

We’re a lean, talent-dense team of explorers, researchers, and full-stack engineers. We move fast, sweat the details, and operate at frontier scale with a roadmap to build the world’s most powerful AI models. Most importantly, we’re united by the belief that doing this right is the only way to do it at all. If our mission resonates with you, we’d love to talk.

Explore all jobs

Related Stories

Hill-climbing MAI models for GitHub Copilot and Excel

July 23, 2026
Models
Superintelligence team

Better models, fewer parameters, less tokens

At Build in June, we introduced our hillclimbing machine, our integrated data, model, and harness flywheel. Today we are excited to share two examples inside Microsoft: MAI models specialized for agentic workloads in GitHub Copilot and Excel.

Early results are promising. In our live product deployment, we see that our MAI model deployed in Excel is on par with GPT-5.6 for the most common tasks while being more cost-efficient.

The figure below shows how MAI-Code-1-Flash, post-trained within the GitHub Copilot harness, was used as the starting checkpoint to climb on Excel evaluations, resulting in two highly efficient, specialized models.

Line graph showing pass rates on SWE Bench (VS Code and Base) across model checkpoints, with data points labeled "Code" and "Excel." Pass rates increase over checkpoints, starting below 72% and peaking at 86%.

MAI-Code-1-Flash in GitHub Copilot

Since launching MAI-Code-1-Flash in GitHub Copilot in June, millions of developers have been using it for their day-to-day work, where it’s outperforming other similarly sized models while using fewer tokens.

  • It has an approximately 10% higher code accept rate than GPT 5.4 Mini and Claude Haiku 4.5 in VS Code.
  • Developers were 6% more likely to return across multiple days than with GPT 5.4 Mini and 11% more likely than with Claude Haiku 4.5.
  • It has 10% lower median token usage than GPT-5.4 mini and Claude Haiku 4.5, with more user-initiated turns.

MAI model live in Excel

Excel offered a test for whether the capabilities built into MAI-Code-1-Flash could transfer beyond the domain they were trained for, moving from agentic coding to agentic knowledge work. To do so, we further trained our MAI-Code-1-Flash checkpoint in an Excel reinforcement learning environment to learn about tools and knowledge workflows in spreadsheets. The result is a model with a command of Excel workflows that is more efficient and less expensive to run.

Flowchart titled "The Excel climb" showing an Excel RL Environment with action, review, execute, and update steps, linked to input, output, grader, and updated model weights.

User feedback from production traffic indicates that the quality of the MAI model in Excel is on par with GPT-5.6 for the most common tasks. In addition to the direct model cost savings, this smaller, more efficient model can be served on both Nvidia H100 and A100 class GPUs rather than requiring only the latest-generation accelerators, which significantly lowers the cost of deployment for Microsoft.

Training agentic models from inside the product stack

These results point toward a broader strategy. By having access to the entire product stack—the model, the harness that runs it, the agents, and product-specific evaluations—we can hill-climb to train efficient, powerful models capable of tasks previously handled by larger, more expensive ones.

Beyond GitHub Copilot and Excel, we’re currently extending this hill-climbing approach to train efficient models across Microsoft’s family of agentic products: Copilot Chat, Outlook, PowerPoint, and more.

Learn More

Related Stories

Introducing MAI-Image-2.5-Pro and MAI-Voice-2-Flash

July 23, 2026
Models
Superintelligence team
Two overlapping speech bubbles, one pink and one green, with a flower in the center where they meet. The background is light gray.

“MAI-Image-2.5-Pro is a strong leap forward for GenMedia tools. Beyond the impressive image quality, its ability to render text with this kind of accuracy is a real breakthrough. It also understands natural language edits, so creative iteration becomes faster and far more intuitive. Microsoft has firmly established itself among the leaders in generative AI.”

Rob Reilly, Global Chief Creative Officer, WPP

A year ago, we set out to develop purpose-built models in-house at Microsoft AI. Models trained on clean, traceable, enterprise-grade data, without distillation from third-party models, and designed from the ground up to serve the people who use Microsoft products every day.

Today, that work is showing up where it matters in the products you already rely on. The models we previewed at Build are not just topping leaderboards, they are running in production, at scale, efficiently powering experiences for millions of users across a growing portfolio of Microsoft products, including Bing, PowerPoint, OneDrive, Dynamics 365, and Azure.

Introducing two new model variants

Real-world use cases are not one-size-fits-all. A creative studio chasing maximum fidelity has very different needs from a customer service center serving millions of calls daily. That’s why we’re building families of models: to give every product the right balance of quality, speed, and cost.

Today we’re expanding those options:

  • MAI-Image-2.5-Pro is now in public preview. For use cases where quality is top of mind. Hero imagery, detailed editing, precise in-image text rendering. Pro is our highest-fidelity image model to date, priced at $5 per 1M text input tokens, $8 per 1M image input tokens, and $106 per 1M image output tokens.
  • MAI-Voice-2-Flash is now in public preview. First introduced at Build, Flash is built for speed and scale. It’s the fast, efficient path for high-volume voice experiences where responsiveness is everything, all while retaining the natural prosody and high acoustic quality found in MAI-Voice-2. Flash is 2x faster than MAI-Voice-2 and 32% cheaper, priced at $15 per 1M characters.

MAI-Voice-2-Flash and MAI-Image-2.5-Pro sit alongside our existing production models so builders can pick the point on the quality-speed-cost curve that fits their job.

A collage of eight images: perfume ad, lemons with a drink, purple tote bag, dog on a crosswalk, city park, two books about birds, feet on checkered tiles, and a modern bathroom with blue accents.

Purpose-built models for our Microsoft products

Leaderboards are great for benchmarking the quality of our models, but the proof lies in serving users in real product use cases. Microsoft rigorously evaluates all models before deciding on the best one for use in production environments. MAI models are outperforming competitive models for quality, latency and efficiency for more of Microsoft’s product surfaces:

  • Bing Image Creator is now 100% in-house by default. With MAI-Image-2.5‘s new precise image editing capabilities, it is now the default model powering Bing Image Creator end-to-end, for high-quality generation and greater creative control.
White text on a brown background reads: "MAI" in the top left and "MAI-Image-2.5 in Bing Image Creator" in large text at the bottom left.
  • Image generation and editing capabilities available in PowerPoint. MAI-Image-2.5 is now in production in PowerPoint for image-to-image capabilities, reducing GPU costs up to 84% compared with GPT-Image-2.
  • White text on a dark brown background reads: "MAI" in the top left and "MAI-Image-2.5 in PowerPoint" in the bottom left.
  • Image editing in OneDrive. MAI-Image-2.5 is now the default model for key OneDrive production image-editing scenarios. Since rollout, it has increased save rates by 26%, reduced P95 latency by approximately 25%, and delivered 2.5x greater efficiency under medium-utilization production workloads.
  • White text on a dark green background says "MAI" in the top left and "MAI-Image-2.5 in OneDrive" in large letters near the bottom left.
  • Voice in the call center. MAI-Voice-2-Flash now powers Dynamics 365 Contact Center, the enterprise platform for building call center agents used by customers like T-Mobile and EasyJet, bringing our most expressive, natural sounding speech to brand defining conversations while reducing GPU costs up to 89%.
  • MAI-Voice-2-Flash integrated in Azure Voice Live. Customers can quickly build voice agents in Voice Live, powered by MAI‑Voice‑2‑Flash’s natural, expressive voices and low latency. Voice Live gives developers a scalable path to building high‑quality agents that support speech‑to‑speech interactions.
  • Each of these enhancements is a step toward the same goal: Microsoft products, powered by Microsoft models, built to serve the people who use them.

    Built with the experts

    Some of today’s most consequential and challenging fields, such as medicine or software engineering, demand more than a general-purpose model. They demand deep expertise, tight feedback loops, and models that are capable of solving complex, real-world challenges. Microsoft AI partners directly with domain specialists to adapt and refine our models for their unique needs:

    • MAI is partnering with Dragon Copilot, a solution used by 170,000 medical providers, which processed 28 million patient encounters last quarter.MAI-Transcribe-1.5 now supports Dragon Copilot’s multilingual workflow, replacing the previous model with our own best-in-class model across 58 languages. In internal evaluations on multilingual recordings, the model delivers a 50% relative reduction in both transcription and language-identification error rates across most languages, and early research shows promising improvements in the downstream accuracy of medical notes.
    Dark green background with "MAI" in small white text at the top left and "MAI-Transcribe in Dragon Copilot" in large white text at the bottom left.

    None of this is an endpoint. It’s the compounding result of one decision: build models in-house so they can be shaped around the people who use our products, and offer options that deliver customers more choice and better value. That’s what “powered by MAI” means, model by model, product by product. We’re just getting started. In the meantime, get started with the models in Foundry. You can also learn more about MAI-Image-2.5-Pro and MAI-Voice-2-Flash on Microsoft Learn or try them out in the MAI Playground.

    Build the Future With Us

    We’re a lean, fast-moving lab made up of some of the world’s most talented minds. We have an exciting roadmap of compute at MAI, with our next-generation GB200 cluster now operational. And we have an ambitious mission we truly believe in. We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in—come and join us as we work on our next generation of models!

    Explore all jobs

    Related Stories

    Two in-house models in support of our mission

    August 28, 2025
    Models

    At Microsoft AI (MAI) we believe AI should be used to empower every person on the planet. We are creating AI for everyone, a supportive, helpful presence always in the service of humanity. It will be the gateway to a universe of knowledge and a set of capabilities that enable people and organizations to achieve more. Responsible, reliable, filled with personality and expertise, we are focused on creating applied AI as a platform for category defining and deeply trusted products that understand each of our unique needs.

    Since last year, we’ve been focused on building the foundation for this vision, with a world class team and infrastructure. To fully meet our goals, MAI requires purpose-built models. Today, we’re excited to preview the first steps to making this a reality.

    • First, we’re releasing MAI-Voice-1, our first highly expressive and natural speech generation model, which is available in Copilot Daily and Podcasts, and as a brand new Copilot Labs experience to try out here. Voice is the interface of the future for AI companions and MAI-Voice-1 delivers high-fidelity, expressive audio across both single and multi-speaker scenarios.
    • Second, we have begun public testing of MAI-1-preview on LMArena, a popular platform for community model evaluation. This represents MAI’s first foundation model trained end-to-end and offers a glimpse of future offerings inside Copilot. We are actively spinning the flywheel to deliver improved models. We’ll have much more to share in the coming months. Stay tuned!

    We have big ambitions for where we go next. Not only will we pursue further advances here, but we believe that orchestrating a range of specialized models serving different user intents and use cases will unlock immense value. There will be a lot more to come from this team on both fronts in the near future. We’re excited by the work ahead as we aim to deliver leading models and put them into the hands of people globally.

    Try MAI-Voice-1 in Copilot and Copilot Labs

    MAI-Voice-1 is a lightning-fast speech generation model, with an ability to generate a full minute of audio in under a second on a single GPU, making it one of the most efficient speech systems available today.

    MAI-Voice-1 is already powering our Copilot Daily and Podcasts features. We are also launching it in Copilot Labs where you can try our expressive speech and storytelling demos. Imagine creating a “choose your own adventure” story with just a simple prompt, or crafting a bespoke guided meditation to help you sleep. Give it a try!

    On a sunny afternoon, a spirited four-year-old named Jamie approached a grizzled pirate who was lounging by the docks. Arr! What be ye wantin’, wee one? This crew ain’t fer the faint of heart! Jamie’s eyes sparkled with excitement as they replied, I wanna be a pirate! I wanna sail the seas and find treasure! Can I join your crew, please? The pirate scratched his beard, chuckling at the child’s enthusiasm. I ye think ye can handle the salty sea air and the dangers of the deep? Jamie nodded vigorously, determination shining through. I can! I can! I’ll be the best pirate ever! The pirate leaned closer, intrigued by Jamie’s spirit. All right, but ye must prove your worth. What be our first task, young matey?

    Under a sprawling Texas sky, a skeptical cowboy and an enthusiastic techie met outside a diner. I reckon this fancy AI voice model ain’t all it’s cracked up to be. Ain’t no machine gonna sound like a real human, the techie chuckled, shaking his head. Oh, come on, this thing can express emotions better than some folks I know. It’s like having a storyteller right in your pocket. The cowboy squinted, pondering the implications of such technology. Maybe so, but can it spin a yarn around a campfire? I ain’t convinced just yet. The techie grinned, undeterred by the cowboy’s skepticism. Just wait till you hear it. It might just surprise you, partner.

    Try MAI-1-preview in LMArena

    MAI-1-preview is an in-house mixture-of-experts model, pre-trained and post-trained on ~15,000 NVIDIA H100 GPUs. This model is designed to provide powerful capabilities to consumers seeking to benefit from models that specialize in following instructions and providing helpful responses to everyday queries.

    We will be rolling MAI-1-preview out for certain text use cases within Copilot over the coming weeks to learn and improve from user feedback. We will continue to use the very best models from our team, our partners, and the latest innovations from the open-source community to power our products. This approach gives us the flexibility to deliver the best outcomes across millions of unique interactions every day.

    In addition to LMArena, we are also making this model available to trusted testers – apply for API access here. We’re excited to collect early feedback to learn more about where the model performs well and how we can make it better. Stay tuned for more.

    Build the future with us

    We’re a lean, fast-moving lab made up of some of the world’s most talented minds. We have an exciting roadmap of compute at MAI, with our next-generation GB200 cluster now operational. And we have an ambitious mission we truly believe in. We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in – come and join us as we work on our next generation of models!

    Explore all jobs

    Related Stories

    English (United States)
    Your Privacy Choices Opt-Out Icon Your Privacy Choices
    Consumer Health Privacy Sitemap Contact Microsoft Privacy Manage cookies Terms of use Trademarks Safety & eco Recycling About our ads