MAI-Voice-2

Turn text into expressive, natural-sounding speech in seconds.

Features

MAI-Voice-2 produces natural, expressive speech from text or a short reference clip, with built-in guardrails ensuring only authorized, consented voices can be used.

Realistic expression

Organic pacing, tone, and emotional range that sound like a person, not a text-to-speech engine.

Voice

Acacia

Joy

Acacia

Anger

Acacia

Disgust

Acacia

Fear

Acacia

Sadness

Emotion

Elm

Joy

Elm

Anger

Elm

Disgust

Elm

Fear

Elm

Sadness

Emotion

Birch

Joy

Birch

Anger

Birch

Disgust

Birch

Fear

Birch

Sadness

Emotion

Grove

Joy

Grove

Anger

Grove

Disgust

Grove

Fear

Grove

Sadness

Emotion

Instant voice matching

Capture any voice from a short reference clip, no fine-tuning needed.
Stable, high-fidelity output that preserves speaker consistency across audiobooks, podcasts, and lectures.
Lectures
Audiobooks
Podcasts
Courses
Documentaries

Natural and expressive across 15 languages

Fluid, emotionally rich speech in 15 languages, without sacrificing quality.

English (US)

Deutsch

Spanish

Français

हिन्दी

Indonesian

Italiano

한국어

Nederlands (Dutch)

Português (Portugal)

Русский

ไทย

Türkçe

Vietnamese

简体中文

Español (México)

Português (Brasil)

Română

Magyar

A blurred image of green foliage and yellow sunlight streaks, creating an abstract, painterly effect against a blue sky background.

Using the Model

Expressive text-to-speech—live and on-demand.

Voice samples generated with MAI-Voice-2 and MAI-Voice-2-Flash

Customer Support

Call-center or customer support: Show a natural back-and-forth where the voice responds quickly enough that it doesn’t feel like a bot “waiting to process.” Interruption handling (if supported) is a great beat here — someone talking over the agent and it adjusting naturally is very persuasive on camera.

MAI-Voice-2-Flash

Shakespearean Wisdom

Behold the silver wanderer of the reeds, gliding soft upon the mirrored dark. With patient poise it waits between the worlds of water, wind, and whispered evening light. A creature not in haste, yet never still, teaching us grace through every careful step.

MAI-Voice-2

Sports Commentator

With everything on the line, the egret makes its move! Slow through the shallows… watching… waiting… And it’s a sudden strike! Got it! Incredible precision from the long beak! The fish never saw it coming. What a scene! Complete composure under pressure. A masterclass performance here in the pond tonight.

MAI-Voice-2

MAI-Voice-2

  • Latency (model‑inference for generating 45s audio)

    1s

  • Price

    $22 per 1M characters

  • Languages

    15+ Languages

    • English (US)
    • English (Australia)
    • Italian
    • French
    • German
    • Hindi
    • Spanish (Spain)
    • Spanish (Mexico)
    • Portuguese (Brazil)
    • Portuguese (Portugal)
    • Korean
    • Chinese (Simplified)
    • Turkish
    • Russian
    • Thai
    • Dutch
    • Romanian
    • Hungarian
  • Granular Emotion Control

    Yes

  • Zero-shot Voice Prompting

    Yes

  • Best For

    Fidelity matters more than speed

    • Audiobooks
    • Content Creation
    • Voice-over
Try in Playground

MAI-Voice-2-Flash

  • Latency (model‑inference for generating 45s audio)

    225ms

  • Price

    $15 per 1M characters

  • Languages

    15+ Languages

    • English (US)
    • English (Australia)
    • Italian
    • French
    • German
    • Hindi
    • Spanish (Spain)
    • Spanish (Mexico)
    • Portuguese (Brazil)
    • Portuguese (Portugal)
    • Korean
    • Chinese (Simplified)
    • Turkish
    • Russian
    • Thai
    • Dutch
    • Romanian
    • Hungarian
  • Granular Emotion Control

    Yes

  • Zero-shot Voice Prompting

    Yes

  • Best For

    Latency sensitive use cases

    • Call Center Agents
    • Voice Assistant
    • IVR
Try in Playground
Performance

Leading in expressiveness and naturalness

MAI-Voice-2 delivers expressive real-time and long-form generation, with stable output and low latency.

Listen across languages

Joy Example

00:00 00:00

Sadness Example

00:00 00:00

Joy Example

00:00 00:00

Sadness Example

00:00 00:00

Joy Example

00:00 00:00

Sadness Example

00:00 00:00

Joy Example

00:00 00:00

Sadness Example

00:00 00:00

Joy Example

00:00 00:00

Sadness Example

00:00 00:00

Joy Example

00:00 00:00

Sadness Example

00:00 00:00

Joy Example

00:00 00:00

Sadness Example

00:00 00:00

Joy Example

00:00 00:00

Sadness Example

00:00 00:00

Joy Example

00:00 00:00

Sadness Example

00:00 00:00

Featured Partner

An older man wearing glasses and a blazer reads a book inside a cozy bookstore filled with shelves and stacks of books.
A stylized logo featuring the letters "P" and "H" intertwined, enclosed within a simple circular border, all in brown on a light background.
“One of [the researchers] recorded my introduction and the next thing I knew, he was playing my voice…and the intonation, the pauses…I just thought, wow, that’s quite nice.
 
“It’s really exciting for me because what we’re about to embark on together is in a way my lifetime’s ambition, which is to bring poetry to everyone.”
– William Sieghart, Ode Founder/Director and author of the “Poetry Pharmacy” anthologies

Try MAI-Voice-2

MAI Playground

Experiment with all other MAI models.
Try in Playground

Copilot Audio Expressions

Bring expressive voice directly into your Copilot workflows.
Try in Copilot

Microsoft Foundry (Azure Speech)

Build and deploy MAI-Voice with Azure Speech.
Try in Azure Speech
English (United States)
Your Privacy Choices Opt-Out Icon Your Privacy Choices
Consumer Health Privacy Sitemap Contact Microsoft Privacy Manage cookies Terms of use Trademarks Safety & eco Recycling About our ads