• About Us
King of Computer Media
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
King of Computer Media
No Result
View All Result

Home - AI Trends and Related News - OpenAI launches three real-time audio models for next-gen voice agents

OpenAI launches three real-time audio models for next-gen voice agents

Claire by Claire
May 8, 2026 - Updated on August 5, 2026
in AI Trends and Related News

OpenAI recently officially released three brand-new real-time audio modelsProviding developers with more powerful tools for building voice applications and intelligent agent systems. These three models are GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper. Their core objectives are to enhance the naturalness of voice interactions, accelerate translation speed, and reduce speech-to-text latency, bringing voice technology closer to the way humans communicate.

OpenAI launches three real-time audio models to power next-generation voice agents

GPT-Realtime-2: Core Flagship Model

Among the three models, GPT-Realtime-2 is considered the most significant breakthrough. It’s specifically designed for real-time voice interaction, capable of understanding user requests, calling external tools, handling voice corrections, and continuing conversations naturally. This enables voice agents to move beyond being one-way command executors, allowing them to respond as flexibly as human assistants.

GPT-Realtime-2 brings multiple new features:

  • Pre-task Response: The model can respond with a brief phrase before executing a task, such as “Let me check,” to make the interaction feel more natural and human-like.
  • Parallel tool calling: It can launch multiple tools simultaneously and report progress to users in real time.
  • Better resilience: when encountering errors or problems, the model doesn’t fail silently, but responds in a more graceful manner.
  • Extended context: The context window has been expanded from 32K to 128K, enabling it to handle more complex contexts.
  • Professional Domain Understanding: The model can more accurately preserve the original meaning in medical, technical, or professional terminology.
  • Tone Control: Adjust the tone based on context to better suit the occasion.
  • Adjustable reasoning difficulty: Developers can choose from different levels of reasoning depth, from minimal to extreme, to optimize performance based on their needs.

In benchmarks, GPT-Realtime-2 significantly outperforms its predecessor. The high-reasoning version achieved a score of 96.6% on the Big Bench Audio test, compared to GPT-Realtime-1.5’s 81.4%. On the Audio MultiChallenge instruction-following test, GPT-Realtime-2’s extreme-reasoning version scored 48.5%, also far exceeding its predecessor’s 34.7%.

GPT-Realtime-Translate: Real-time Multilingual Translation

Another new model, GPT-Realtime-Translate, focuses on real-time multilingual voice experiences, translating over 70 input languages into 13 output languages. OpenAI emphasizes that even when users switch contexts mid-conversation, use regional accents, or employ specialized terminology, the model maintains speech speed while accurately preserving meaning. This means cross-language communication will be smoother, benefiting scenarios ranging from international conferences and cross-border customer service to multilingual learning environments.

GPT-Realtime-Whisper: Low-Latency Speech-to-Text

The third model, GPT-Realtime-Whisper, is a streaming transcription system designed specifically for low latency. It can transcribe audio in real-time while someone is speaking, making it suitable for scenarios such as live captions, meeting notes, and classroom notes. The value of this technology lies in the fact that it not only improves efficiency but also enables faster conversion of voice data into text, facilitating subsequent searching, organization, and analysis.

Pricing Standards and Use Cases

These three models are now available through OpenAI’s Realtime API, with clear pricing:

  • GPT-Realtime-2: Audio input tokens at $32 per million, cached input tokens at $0.40 per million, and audio output tokens at $64 per million.
  • GPT-Realtime-Translate: $0.034 per minute.
  • GPT-Realtime-Whisper: $0.017 per minute.

Developers can Playground Try these models directly to quickly experience their capabilities; for general users, OpenAI continues working to integrate these technologies into ChatGPT’s voice experience, enabling more people to enjoy natural, real-time voice interactions.

Source: KOCPC Chinese

Tags: aiGPTModelOPENAIPlease provide the Traditional Chinese text you would like me to translate to English.Transcriptionvoice model

Recent Posts

  • The Xiaomi Pad 8S Pro has passed network access certification and will debut with the self-developed XRING O3 chip.
  • The entire Google Pixel 11 lineup has been leaked! Official promotional renders of the Pixel 11 Pro XL have also surfaced
  • Are Chinese phone battery capacities falsely labeled? A brief look at the “capacity locking” phenomenon in Chinese silicon-carbon batteries.
  • NCC is leaderless, recklessly sending out national-level alert messages!?
  • What does “QR” in QR Code mean?

Recent Comments

No comments to show.
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology

No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology