• About Us
King of Computer Media
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
King of Computer Media
No Result
View All Result

Home - AI Trends and Related News - Microsoft Rolls Out Three MAI Models at Once: Transcribe-1 Dominates Speech-to-Text, Voice-1 Generates Speech at Lightning Speed, Image-2 Ranks Among the Top Three

Microsoft Rolls Out Three MAI Models at Once: Transcribe-1 Dominates Speech-to-Text, Voice-1 Generates Speech at Lightning Speed, Image-2 Ranks Among the Top Three

KOCPC Editor by KOCPC Editor
April 5, 2026 - Updated on August 5, 2026
in AI Trends and Related News, Latest Technology News

Recently, Microsoft, which has been quite aggressive in the AI spaceReleased three new proprietary models for speech-to-text, speech generation, and image generation all at once: specialized in speech-to-text MAI-Transcribe-1Natural speech synthesis MAI-Voice-1and significant advances in image generation capabilities MAI-Image-2And it will be simultaneously available to developers through the Microsoft Foundry platform starting today, with highly competitive pricing. This is also the first major achievement showcased by Microsoft’s AI Super Intelligence team since its establishment six months ago. The team is led by Mustafa Suleyman, a former Google AI executive who joined Microsoft in 2025.

Microsoft releases three MAI models at once

MAI-Transcribe-1: Leading the Industry with 25 Languages, Only 3.9% Error Rate

MAI-Transcribe-1 is the highlight of this release. This speech-to-text model achieves an average Word Error Rate (WER) of only in the FLEURS multilingual speech benchmark, an industry standard, across the top 25 most used languages globally (including Japanese). 3.9%, the lowest record among all models in this benchmark. Microsoft also released detailed comparison data with its competitors, with MAI-Transcribe-1 comprehensively outperforming OpenAI’s Whisper-large-v3 across these 25 languages, among which 22 languagesOutperformed Google’s Gemini 3.1 Flash, and even beat ElevenLabs’ Scribe v2 and OpenAI’s GPT-Transcribe on 15 different languages.

In addition to leading in accuracy, MAI-Transcribe-1 also claims faster processing speeds compared to Azure’s existing Fast speech service 2.5 timesThe model supports mainstream audio formats such as MP3, WAV, and FLAC, with a maximum file size of 200MB. Features like speaker diarization, context biasing, and streaming are currently listed as “coming soon.” For commercial applications, MAI-Transcribe-1 is priced at per hour of audio. $0.36(approximately NT$12), an exceptional value

MAI-Voice-1: Generate 60 seconds of natural speech in one second

MAI-Voice-1 is Microsoft’s latest text-to-speech (TTS) model, designed to generate natural and expressive speech. According to the official description, MAI-Voice-1 can 1 secondEndogenous growth reaches 60 secondsdelivers high-quality voice content with exceptional GPU utilization efficiency, achieving an excellent balance between generation speed and cost.

A major selling point of this model is its ability to maintain the speaker’s voice characteristics and personal traits consistently over extended periods. Developers only need to provide a few seconds of voice samples to create exclusive custom voice profiles through Microsoft Foundry. MAI-Voice-1 is currently applied in the Copilot Audio Expressions feature. The pricing is per 100,000 characters for $22(approximately NT$720)

MAI-Image-2: Image Generation Ranks Among Arena.ai’s Top Three

MAI-Image-2 is not an entirely new release: it was quietly launched on March 19, 2026, and is now officially incorporated into the MAI model family with full commercial availability. This image generation model has already been recognized by industry-renowned Arena.ai ranks in the top three on image generation leaderboardsSince its launch, it has continued to power Copilot’s image generation capabilities. Microsoft states that MAI-Image-2 can generate natural lighting effects, precise skin tones and textures, with charts, layouts, and even text within images displayed clearly. Compared to the previous version, MAI-Image-2 delivers the same quality output on the Foundry and Copilot platforms with improved generation speed. At least 2 timesOne of the world’s largest advertising groups WPP Already pioneering as an enterprise partner, adopting MAI-Image-2 at scale for commercial creation.

MAI-Image-2 pricing: text input per 1 million tokens for $5(approximately NT$165), image generation each 1 million tokens for $33(approximately NT$1,080)

Suleiman’s First Interview: Breaking Free from Six Years Under OpenAI’s Exclusive Contract

The significance of this three-pronged release extends far beyond the model itself, as foreign media…Exclusive interviewRevealed the real reason behind Microsoft: the contract between Microsoft and OpenAI that began in 2019 did not undergo major revision until September 2025. According to Suleiman, the old contract explicitly prohibited Microsoft from independently developing Artificial General Intelligence (AGI) or superintelligence before 2032, which long constrained Microsoft to OpenAI’s technical approach in frontier AI model development. The new contract in October 2025 finally freed Microsoft from this constraint.

Suleiman said: “In September last year, we renegotiated the contract with OpenAI, which allows us to independently pursue our own superintelligence research. Since then, we have assembled the necessary compute, built the team, and acquired the necessary data.” He also emphasized that the partnership with OpenAI remains unchanged and will continue at least until 2032.

Suleiman further revealed his pride about this release: “The model we’re releasing today is the world’s top in speech-to-text: not only that, we can deliver the same level of quality using half the GPU computing resources of our competitors.” The backdrop to these remarks: Microsoft’s stock price just posted its worst quarterly performance since the 2008 financial crisis, with investor doubts about whether massive AI infrastructure investments can translate into substantial revenue growing louder. These three MAI models can be seen as Suleiman and Microsoft’s leadership’s first formal response to Wall Street.

Source

Source: KOCPC Chinese

Tags: aiMAIMAI-Image-2MAI-Transcribe-1MAI-Voice-1Microsoft

Recent Posts

  • The Xiaomi Pad 8S Pro has passed network access certification and will debut with the self-developed XRING O3 chip.
  • The entire Google Pixel 11 lineup has been leaked! Official promotional renders of the Pixel 11 Pro XL have also surfaced
  • Are Chinese phone battery capacities falsely labeled? A brief look at the “capacity locking” phenomenon in Chinese silicon-carbon batteries.
  • NCC is leaderless, recklessly sending out national-level alert messages!?
  • What does “QR” in QR Code mean?

Recent Comments

No comments to show.
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology

No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology