According to a well-known AI developer on the X platform, OpenAI is preparing to launch a new generation speech model codenamed GPT-Bidi-1, which will be the largest upgrade in the history of ChatGPT speech model. According to a report by technology media testingcatalog, GPT-Bidi-1 adopts a “two-way” (BiDi) architecture, which can listen and speak at the same time, absorb user interruptions during the conversation, and adjust response content in real time. This upgrade is believed to completely change the smoothness and naturalness of ChatGPT’s voice conversations, making AI voice conversations close to the level of real-person interaction for the first time. It will also make ChatGPT more competitive in the voice assistant market and compete head-on with giants such as Google and Apple.

OpenAI prepares to launch GPT-Bidi-1 bidirectional speech model
This news was first exposed by developer M1Astra on The developer community has responded enthusiastically to this, and many believe this is the most important update to ChatGPT’s voice capabilities. It will directly challenge the voice conversation capabilities of Google Gemini Live and Apple Siri, and inject new variables into the AI voice assistant market.
New OpenAI voice model “GPT-Bidi-1”
Coming soon with a “major leap in intelligence”
– The next generation of Voice
– More natural conversations, powered by our next-generation voice model https://t.co/mvH9TSisgO pic.twitter.com/Ka3Mk2LpXV— M1 (@M1Astra) June 16, 2026
Two-way architecture: a major breakthrough in voice dialogue
The core technology of GPT-Bidi-1 is the “BiDi” architecture, which is a new generation of speech processing method developed by OpenAI starting in early 2026. The biggest difference from existing speech models is that the BiDi model can process auditory input and speech output at the same time, instead of having to wait for the user to finish speaking before starting to respond like the current system. This technology was first reported by The Information in early 2026. At that time, OpenAI internally positioned it as the next-generation core architecture of voice technology, with the goal of making AI voice conversations as smooth and natural as human conversations, achieving true two-way communication.

This means that users can interrupt ChatGPT while it is speaking, correct its direction, or send out confirmation signals such as “um hum”, and the model will sense and adjust its response in real time. In contrast, the current ChatGPT voice mode will lock the response once it starts speaking, and the user’s interruption in the middle will often cause the conversation to get stuck and need to be restarted. The BiDi architecture makes conversations more like real human communication rather than machines taking turns, greatly reducing the stiffness of voice conversations and making the overall user experience more intuitive and natural.
According to OpenAI internal information, after GPT-Bidi-1 is launched, users can switch between two-way mode and the existing advanced voice mode according to needs, and supports three intelligence levels: High, Medium and Instant, allowing users to adjust the response speed and depth according to task requirements. For example, when you need to quickly check the weather or time, you can choose Instant mode to get an immediate response. When having complex discussions or brainstorming, you can switch to High mode to get deeper and more insightful answers.
Lagging behind and catching up in voice technology
Currently, OpenAI’s text model has rapidly evolved to the GPT-5.5 generation, but the voice function still remains on the older audio technology route, resulting in spoken dialogue capabilities significantly lagging behind text performance. The launch of GPT-Bidi-1 is to bridge this gap and make ChatGPT’s voice conversation capabilities catch up with text levels. This also means that ChatGPT’s voice mode will adopt the same intelligence level as the text model for the first time, and will no longer be just a “speaking text model”.
This gap is critical to OpenAI, which is betting that speech, rather than text, will become the primary way humans interact with AI. The AI smart speaker developed by OpenAI in cooperation with former Apple design chief Jony Ive (expected to sell for US$200 to US$300, approximately NT$6,500 to NT$9,750, and be launched as soon as February 2027) requires a two-way speech engine like GPT-Bidi-1 as the core. On a device without a screen, the ability to speak naturally isn’t a plus, it’s the entire interface. This also explains why OpenAI attaches so much importance to the upgrade of voice technology and invests a lot of resources to catch up with Google and Apple’s leadership in the voice field to ensure that future hardware products are sufficiently competitive.
Development history and time to market
OpenAI’s BiDi speech model was originally targeted to be launched in the first quarter of 2026, but it is rumored that the prototype will have problems such as abnormal sounds after several minutes of continuous conversation, causing the schedule to be delayed to the second quarter or later. Signs of preparation for GPT-Bidi-1 have now appeared in the ChatGPT web and mobile versions, indicating that the release of the consumer side is close, but the final naming may still be adjusted. It is expected that GPT-Bidi-1 may be officially unveiled in the next few weeks. It will compete head-on with competing products such as Google’s Gemini Live and Apple’s Siri, setting off a new wave of AI voice wars and redefining the standards for AI voice conversations.
GPT-Bidi-1 is also considered to significantly improve the application capabilities of ChatGPT in customer service scenarios. The BiDi architecture is able to call external tools and applications while speaking, which is a key requirement for customer service scenarios that require real-time querying of databases or performing operations. OpenAI has always regarded customer service automation as an important application scenario of voice technology. GPT-Bidi-1’s two-way dialogue capability allows AI customer service to handle customers’ questions and needs more naturally, while querying order information or performing refund operations in the background, making the service process smoother.
Summarize
The launch of GPT-Bidi-1 will be the biggest upgrade to ChatGPT voice mode since its launch. The shift from “taking turns to speak” to “simultaneous listening and speaking” may seem like just an adjustment of technical details, but in fact it will completely change the experience of human interaction with AI voice. For OpenAI, which is building screen-less AI hardware, this technology is an indispensable key puzzle. With the imminent release of GPT-Bidi-1, ChatGPT’s voice dialogue capabilities are expected to undergo a real qualitative change, making the dialogue between users and AI more natural and smooth, and laying a more solid foundation for OpenAI’s voice ecosystem and future hardware products.
Source: KOCPC Chinese