On October 1, OpenAI held its developer conference (OpenAI DevDay) in San Francisco, USA. This OpenAI DevDay introduced innovative features such as the Realtime API, prompt caching, and vision fine-tuning. Developers can now use OpenAI’s advanced voice mode in apps for a low-latency multimodal voice experience. The conference also shared examples of partners using GPT-4o. Let’s take a look at the new features from this 2024 OpenAI DevDay!

OpenAI Developer Conference New Feature Roundup: Real-Time API, Prompt Caching, Vision Fine-Tuning
At the OpenAI DevDay on October 1st, many new features were announced, and among them, the one the editor found most exciting wasRealtime APIOpenAI launches the public beta of its real-time API, now allowing all paid developers to build low-latency, multimodal voice experiences similar to ChatGPT’s advanced voice mode within their apps.

OpenAI has already been testing the real-time API with a small number of partners. For instance, nutrition and fitness coaching app Healthify and language learning app Speak are already using the real-time API to support their work.
What an unforgettable experience at OpenAI’s DevDay summit! 🚀 We showcased the power of the new Realtime API, demonstrating how Healthify integrated it to create low-latency, multimodal conversational experiences, bringing personalized health coaching to life like never before.… pic.twitter.com/sBJAb0y7bQ
— Healthify (@healthifyme) October 3, 2024
1/ I’m excited for Speak to be a featured use case for @OpenAI’s Realtime API launch today! We’ve been working closely with them on this API for a while. And speech-to-speech is going to be a game changer not just for Speak but for all software.https://t.co/V7QtYa61Xb pic.twitter.com/CQhFBRqzOw
— Connor Zwick (@connorzwick) October 1, 2024
The real-time API can also be combined with other apps’ APIs; for example, OpenAI demonstrated how the real-time API can work together with Twilio’s API. An AI assistant calls a fictional candy store over the phone to order 400 chocolate strawberries, and Twilio’s API then consolidates the order details—such as the items ordered, order contents, delivery location, preparation time, and so on—giving Twilio the ability to handle phone communication tasks.
OpenAI beat Google to real time phone calls!? pic.twitter.com/uU6ELIp2ld
— Nick Dobos (@NickADobos) October 1, 2024
Cache Prompt (Prompt Caching in the API)
OpenAI also launched this timeCache Prompt feature, allowing developers to cache frequently used context messages across multiple API requests. This helps reduce developers’ costs and time. According to OpenAI, the prompt caching feature can save developers up to about 50% in expenses.

Vision Fine-Tuning
OpenAI launches on GPT-4oVision Fine-Tuning featureBesides text, images can now also be used for fine-tuning. Developers can customize the model to give GPT-4o stronger image understanding capabilities, enabling better development in fields such as autonomous driving, video detection, medical image analysis, and visual search in the future.

Grab, which provides food delivery and ride-hailing services in Southeast Asia, uses GPT-4o’s vision fine-tuning feature to convert the street-view images collected by their drivers into map data, which powers its own GrabMaps. Grab has now taught GPT-4o to correctly locate traffic signs and count lane dividers, and as a result, Grab has improved lane count accuracy by 20% and speed limit sign positioning accuracy by 13%.

This concludes the introduction to the OpenAI developer conference. According to OpenAI, in addition to the developer conference on October 1, there will be two more launch events in London and Singapore on October 30 and November 21. I wonder what new features OpenAI will introduce then. If you’re interested in the OpenAI San Francisco developer conference, you canGo to the OpenAI official website.Learn more:

Source: KOCPC Chinese