Developers can now integrate a full duplex voice system into applications that distinguishes speech from background noise and supports natural interruptions. The model combines listening and speaking in a single architecture to eliminate handoffs between separate audio and text processes. Tool calling for spoken requests has an 87% pass rate, allowing AI agents to delegate complex reasoning to backend models such as GPT-6 Astra.

The new API costs $0.05 per minute for the voice layer and lowers turn taking latency to 0.798 seconds from 1.41 seconds for GPT-Realtime-2.1. Based on the technology provided to 1 billion ChatGPT users, the system achieves an 86.2% score on Tau3 voice tasks compared to 45.7% for the previous version. Early deployments include restaurant service tools from Yelp Host and Hatch, and language lessons from Speak, which saw nearly 80% fewer interruptions in early testing.

Sign in to suggest edits

Key sources

  1. SOURCE@openaidevs“voice agents that listen while they speak”x.com
  2. SUPPORT@openaidevs“handles listening and speaking in one model”x.com
  3. SUPPORT@openaidevs“distinguishes speech from background noise, so café chatter doesn’t have to stop the conversation”x.com
  4. SUPPORT@juberti“less expensive - only $.05/minute!”x.com
  5. SUPPORT@thsottiaux“same voice system that we shipped to 1B users in ChatGPT”x.com
  6. SUPPORT@wallstengine“86.2% on Tau3 voice tasks when paired with Astra, vs. 45.7% for GPT-Realtime-2.1”x.com
  7. SUPPORT@openaidevs“Yelp Host & @usehatchapp are among the first to deploy @OpenAI's GPT-Live-1”x.com
  8. DISCUSSIONarittrnews.ycombinator.com
Markdown