Sand dunes covered in tall grass under a clear blue sky, with sunlight casting shadows on the landscape.

MAI-Voice-2.1

Turn text into expressive, natural-sounding speech in seconds.

Features

With the release of MAI-Voice-2.1 and Voice-2.1-Flash, developers now have the best of both worlds: speed and quality.

Together, they improve both sides of a voice conversation, from understanding speech in real time to generating natural responses faster.

Instant voice matching

Capture any voice from a short reference clip, no fine-tuning needed.
A blurred image of green foliage and yellow sunlight streaks, creating an abstract, painterly effect against a blue sky background.

Using the Model

Expressive text-to-speech. Live and on-demand.

Voice samples generated with MAI-Voice-2.1 and MAI-Voice-2.1-Flash

Customer Support

Call-center or customer support: Show a natural back-and-forth where the voice responds quickly enough that it doesn’t feel like a bot “waiting to process.” Interruption handling (if supported) is a great beat here — someone talking over the agent and it adjusting naturally is very persuasive on camera.

MAI-Voice-2.1-Flash

Shakespearean Wisdom

O fair barista, lend thine ear to me. In morning’s light, I crave a potion warm, a latte rich, with froth like clouds above. Sweet nectar of the bean, I do implore.

MAI-Voice-2.1

Sports Commentator

We’re deep into stoppage time, and the atmosphere is electric! González whips in a perfect cross, Martínez rises, hangs in the air like a superhero! What a header! That’s pure magic from the number 9!

MAI-Voice-2.1

Meditation

Find a cozy spot, my dear, whether you choose to sit, lie down, or simply stand in serene stillness. Gently close your eyes if that feels right for you. Now take a slow, deep breath in, filling your lungs, and exhale fully, releasing the weight of the day.

MAI-Voice-2.1

Pirate

Yarr! Gather round, ye skellywags. It be I, Cap’n Barnacle Bill, the unluckiest pirate to ever roam the seven seas.

MAI-Voice-2.1

Performance

Leading in expressiveness and naturalness

MAI-Voice-2.1 delivers expressive real-time and long-form generation, with stable output and low latency.

MAI-Voice-2.1

  • Latency (model inference)

    ~550ms

  • Price

    $22 per 1M characters

  • Languages

    23 Languages

    • English (US)
    • English (Australia)
    • English (UK)
    • English (India)
    • Italian
    • French
    • German
    • Hindi
    • Spanish (Spain)
    • Spanish (Mexico)
    • Portuguese (Brazil)
    • Portuguese (Portugal)
    • Korean
    • Chinese (Simplified)
    • Turkish
    • Russian
    • Thai
    • Dutch
    • Romanian
    • Hungarian
    • Czech
    • Danish
    • Finnish
    • Indonesian
    • Polish
    • Swedish
    • Norwegian (Bokmål)
    • Vietnamese
  • Granular Emotion Control

    Yes

  • Zero-shot Voice Prompting

    Yes

  • Best For

    Fidelity matters more than speed

    • Audiobooks
    • Content Creation
    • Voice-over
Try in Playground

MAI-Voice-2.1-Flash

  • Latency (model inference)

    ~45ms

  • Price

    $15 per 1M characters

  • Languages

    23 Languages

    • English (US)
    • English (Australia)
    • English (UK)
    • English (India)
    • Italian
    • French
    • German
    • Hindi
    • Spanish (Spain)
    • Spanish (Mexico)
    • Portuguese (Brazil)
    • Portuguese (Portugal)
    • Korean
    • Chinese (Simplified)
    • Turkish
    • Russian
    • Thai
    • Dutch
    • Romanian
    • Hungarian
    • Czech
    • Danish
    • Finnish
    • Indonesian
    • Polish
    • Swedish
    • Norwegian (Bokmål)
    • Vietnamese
  • Granular Emotion Control

    Yes

  • Zero-shot Voice Prompting

    Yes

  • Best For

    Latency sensitive use cases

    • Call Center Agents
    • Voice Assistant
    • IVR
Try in Playground

Featured Partner

An older man wearing glasses and a blazer reads a book inside a cozy bookstore filled with shelves and stacks of books.
A stylized logo featuring the letters "P" and "H" intertwined, enclosed within a simple circular border, all in brown on a light background.
“One of [the researchers] recorded my introduction and the next thing I knew, he was playing my voice…and the intonation, the pauses…I just thought, wow, that’s quite nice.
 
“It’s really exciting for me because what we’re about to embark on together is in a way my lifetime’s ambition, which is to bring poetry to everyone.”
– William Sieghart, Ode Founder/Director and author of the “Poetry Pharmacy” anthologies

Try MAI-Voice-2.1

MAI Playground

Experiment with all other MAI models.
Try in Playground

Copilot Audio Expressions

Bring expressive voice directly into your Copilot workflows.
Try in Copilot

Microsoft Foundry (Azure Speech)

Build and deploy MAI-Voice with Azure Speech.
Try in Azure Speech
English (United States)
Your Privacy Choices Opt-Out Icon Your Privacy Choices
Consumer Health Privacy Sitemap Contact Microsoft Privacy Manage cookies Terms of use Trademarks Safety & eco Recycling About our ads