Introducing MAI-Image-2.5-Pro and MAI-Voice-2-Flash

July 23, 2026
Models
Superintelligence team
Two overlapping speech bubbles, one pink and one green, with a flower in the center where they meet. The background is light gray.

“MAI-Image-2.5-Pro is a strong leap forward for GenMedia tools. Beyond the impressive image quality, its ability to render text with this kind of accuracy is a real breakthrough. It also understands natural language edits, so creative iteration becomes faster and far more intuitive. Microsoft has firmly established itself among the leaders in generative AI.”

Rob Reilly, Global Chief Creative Officer, WPP

A year ago, we set out to develop purpose-built models in-house at Microsoft AI. Models trained on clean, traceable, enterprise-grade data, without distillation from third-party models, and designed from the ground up to serve the people who use Microsoft products every day.

Today, that work is showing up where it matters in the products you already rely on. The models we previewed at Build are not just topping leaderboards, they are running in production, at scale, efficiently powering experiences for millions of users across a growing portfolio of Microsoft products, including Bing, PowerPoint, OneDrive, Dynamics 365, and Azure.

Introducing two new model variants

Real-world use cases are not one-size-fits-all. A creative studio chasing maximum fidelity has very different needs from a customer service center serving millions of calls daily. That’s why we’re building families of models: to give every product the right balance of quality, speed, and cost.

Today we’re expanding those options:

  • MAI-Image-2.5-Pro is now in public preview. For use cases where quality is top of mind. Hero imagery, detailed editing, precise in-image text rendering. Pro is our highest-fidelity image model to date, priced at $5 per 1M text input tokens, $8 per 1M image input tokens, and $106 per 1M image output tokens.
  • MAI-Voice-2-Flash is now in public preview. First introduced at Build, Flash is built for speed and scale. It’s the fast, efficient path for high-volume voice experiences where responsiveness is everything, all while retaining the natural prosody and high acoustic quality found in MAI-Voice-2. Flash is 2x faster than MAI-Voice-2 and 32% cheaper, priced at $15 per 1M characters.

MAI-Voice-2-Flash and MAI-Image-2.5-Pro sit alongside our existing production models so builders can pick the point on the quality-speed-cost curve that fits their job.

A collage of eight images: perfume ad, lemons with a drink, purple tote bag, dog on a crosswalk, city park, two books about birds, feet on checkered tiles, and a modern bathroom with blue accents.

Purpose-built models for our Microsoft products

Leaderboards are great for benchmarking the quality of our models, but the proof lies in serving users in real product use cases. Microsoft rigorously evaluates all models before deciding on the best one for use in production environments. MAI models are outperforming competitive models for quality, latency and efficiency for more of Microsoft’s product surfaces:

  • Bing Image Creator is now 100% in-house by default. With MAI-Image-2.5‘s new precise image editing capabilities, it is now the default model powering Bing Image Creator end-to-end, for high-quality generation and greater creative control.
White text on a brown background reads: "MAI" in the top left and "MAI-Image-2.5 in Bing Image Creator" in large text at the bottom left.
  • Image generation and editing capabilities available in PowerPoint. MAI-Image-2.5 is now in production in PowerPoint for image-to-image capabilities, reducing GPU costs up to 84% compared with GPT-Image-2.
  • White text on a dark brown background reads: "MAI" in the top left and "MAI-Image-2.5 in PowerPoint" in the bottom left.
  • Image editing in OneDrive. MAI-Image-2.5 is now the default model for key OneDrive production image-editing scenarios. Since rollout, it has increased save rates by 26%, reduced P95 latency by approximately 25%, and delivered 2.5x greater efficiency under medium-utilization production workloads.
  • White text on a dark green background says "MAI" in the top left and "MAI-Image-2.5 in OneDrive" in large letters near the bottom left.
  • Voice in the call center. MAI-Voice-2-Flash now powers Dynamics 365 Contact Center, the enterprise platform for building call center agents used by customers like T-Mobile and EasyJet, bringing our most expressive, natural sounding speech to brand defining conversations while reducing GPU costs up to 89%.
  • MAI-Voice-2-Flash integrated in Azure Voice Live. Customers can quickly build voice agents in Voice Live, powered by MAI‑Voice‑2‑Flash’s natural, expressive voices and low latency. Voice Live gives developers a scalable path to building high‑quality agents that support speech‑to‑speech interactions.
  • Each of these enhancements is a step toward the same goal: Microsoft products, powered by Microsoft models, built to serve the people who use them.

    Built with the experts

    Some of today’s most consequential and challenging fields, such as medicine or software engineering, demand more than a general-purpose model. They demand deep expertise, tight feedback loops, and models that are capable of solving complex, real-world challenges. Microsoft AI partners directly with domain specialists to adapt and refine our models for their unique needs:

    • MAI is partnering with Dragon Copilot, a solution used by 170,000 medical providers, which processed 28 million patient encounters last quarter.MAI-Transcribe-1.5 now supports Dragon Copilot’s multilingual workflow, replacing the previous model with our own best-in-class model across 58 languages. In internal evaluations on multilingual recordings, the model delivers a 50% relative reduction in both transcription and language-identification error rates across most languages, and early research shows promising improvements in the downstream accuracy of medical notes.
    Dark green background with "MAI" in small white text at the top left and "MAI-Transcribe in Dragon Copilot" in large white text at the bottom left.

    None of this is an endpoint. It’s the compounding result of one decision: build models in-house so they can be shaped around the people who use our products, and offer options that deliver customers more choice and better value. That’s what “powered by MAI” means, model by model, product by product. We’re just getting started. In the meantime, get started with the models in Foundry. You can also learn more about MAI-Image-2.5-Pro and MAI-Voice-2-Flash on Microsoft Learn or try them out in the MAI Playground.

    Build the Future With Us

    We’re a lean, fast-moving lab made up of some of the world’s most talented minds. We have an exciting roadmap of compute at MAI, with our next-generation GB200 cluster now operational. And we have an ambitious mission we truly believe in. We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in—come and join us as we work on our next generation of models!

    Explore all jobs

    Related Stories

    Hill-Climbing MAI models for GitHub Copilot and Excel

    July 23, 2026
    Models
    Superintelligence team

    Better models, fewer parameters, less tokens

    At Build in June, we introduced our hillclimbing machine, our integrated data, model, and harness flywheel. Today we are excited to share two examples inside Microsoft: MAI models specialized for agentic workloads in GitHub Copilot and Excel.

    Early results are promising. In our live product deployment, we see that our MAI model deployed in Excel is on par with GPT-5.6 for the most common tasks while being more cost-efficient.

    The figure below shows how MAI-Code-1-Flash, post-trained within the GitHub Copilot harness, was used as the starting checkpoint to climb on Excel evaluations, resulting in two highly efficient, specialized models.

    Line graph showing pass rates on SWE Bench (VS Code and Base) across model checkpoints, with data points labeled "Code" and "Excel." Pass rates increase over checkpoints, starting below 72% and peaking at 86%.

    MAI-Code-1-Flash in GitHub Copilot

    Since launching MAI-Code-1-Flash in GitHub Copilot in June, millions of developers have been using it for their day-to-day work, where it’s outperforming other similarly sized models while using fewer tokens.

    • It has an approximately 10% higher code accept rate than GPT 5.4 Mini and Claude Haiku 4.5 in VS Code.
    • Developers were 6% more likely to return across multiple days than with GPT 5.4 Mini and 11% more likely than with Claude Haiku 4.5.
    • It has 10% lower median token usage than GPT-5.4 mini and Claude Haiku 4.5, with more user-initiated turns.

    MAI model live in Excel

    Excel offered a test for whether the capabilities built into MAI-Code-1-Flash could transfer beyond the domain they were trained for, moving from agentic coding to agentic knowledge work. To do so, we further trained our MAI-Code-1-Flash checkpoint in an Excel reinforcement learning environment to learn about tools and knowledge workflows in spreadsheets. The result is a model with a command of Excel workflows that is more efficient and less expensive to run.

    Flowchart titled "The Excel climb" showing an Excel RL Environment with action, review, execute, and update steps, linked to input, output, grader, and updated model weights.

    User feedback from production traffic indicates that the quality of the MAI model in Excel is on par with GPT-5.6 for the most common tasks. In addition to the direct model cost savings, this smaller, more efficient model can be served on both Nvidia H100 and A100 class GPUs rather than requiring only the latest-generation accelerators, which significantly lowers the cost of deployment for Microsoft.

    Training agentic models from inside the product stack

    These results point toward a broader strategy. By having access to the entire product stack—the model, the harness that runs it, the agents, and product-specific evaluations—we can hill-climb to train efficient, powerful models capable of tasks previously handled by larger, more expensive ones.

    Beyond GitHub Copilot and Excel, we’re currently extending this hill-climbing approach to train efficient models across Microsoft’s family of agentic products: Copilot Chat, Outlook, PowerPoint, and more.

    Learn More

    Related Stories

    English (United States)
    Your Privacy Choices Opt-Out Icon Your Privacy Choices
    Consumer Health Privacy Sitemap Contact Microsoft Privacy Manage cookies Terms of use Trademarks Safety & eco Recycling About our ads