• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    OpenAI announces GPT-Live-1 API availability

    OpenAI says GPT-Live-1 supports voice agents that listen while they speak and work with developers’ chosen models and harnesses.

    SB
    DF
    BM
    16 Sources, 20d ago, first seen 20d ago

    TLDR

    OpenAI announced GPT-Live-1’s API availability on September 10, describing it as a way to bring ChatGPT’s natural back-and-forth to other apps. The company says the voice agents can listen while speaking and work with the models and harnesses developers choose.

    Combined views

    970.3K

    16 Sources, first seen 20d ago

    likes

    Combined views

    970.3K

    16 Sources, first seen 20d ago

    4.3K likes
    4.3K
    273 comments
    1.9K saves
    265 reposts
    273 comments
    1.9K saves
    265 reposts

    Sentiment

    Positive87.2%12.8%Negative

    Summary

    Sentiment

    Positive87.2%12.8%Negative

    Positive replies hailed GPT-Live-1’s realtime API and full-duplex voice apps as a massive leap toward human-like conversations, while negative replies called out high hourly costs and problems such as lagging or regional unavailability.

    Based on 132 sentiment-bearing replies from 109 accounts across 8 conversations.

    Summary

    Positive replies hailed GPT-Live-1’s realtime API and full-duplex voice apps as a massive leap toward human-like conversations, while negative replies called out high hourly costs and problems such as lagging or regional unavailability.

    Based on 132 sentiment-bearing replies from 109 accounts across 8 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    16 Sources

    @bnicholehopkinsOpenAI just shipped GPT-Live-1 in the API, and I sat down with Peter Bakkum, Member of the Technical Staff working on the realtime APIs to hear his take on how this changes the game. I have been so excited for this launch - over the past few months we've been working with Open AI in alpha testing the latest speech to speech API and I'm really impressed. Instead of chaining together separate speech-to-text, language, and text-to-speech models, GPT-Live-1 reasons across incoming and outgoing audio as one continuous interaction. That matters because VAD and turn detection have wreaked havoc on voice agents. Awkward pauses, interruptions, dropped context, background noise, and accidental responses to silence can all make an agent feel painfully robotic. GPT-Live-1 is designed around a more fluid, turnless interaction model. The result is a conversation that feels much closer to how people actually speak. It also introduces a compelling architecture for balancing natural conversation with deep intelligence. The voice model can stay fast and context-aware while delegating complex reasoning and tool calls to a backend model—including third-party models. Developers can shape tone, pace, and conversational style directly through the system prompt. Longer sessions retain context and quality more reliably. And full-duplex telephony support opens the door to real phone-based agents for customer support, restaurant reservations, scheduling, and much more. In our conversation, Peter and I cover: - What made this API particularly difficult to build and deploy - How developers can prepare to migrate from cascaded architectures - Why voice evaluations are essential - How full-duplex models change the way we build voice agents - The entirely new categories of applications this technology could unlock Voice is becoming a the interface of AI and this release brings that future meaningfully closer. Listen to the full conversation at the link below. 🎙️ Youtube: https://youtu.be/yr-Em6RL7mM Spotify: https://open.spotify.com/episode/21RZ91AMKuVbeMUrn2gknw?si=3071341125ae4ecb
    @muratcanWhen I tried GPT-Live-1 today, I felt deflated as I recognized behaviors we’d spent weeks trying to make reliable. For the past three months, I’ve been focusing on researching and building duplex voice agents, where I wrote thousands of lines across our harness and the experiments around it. Customers use parts of that work every day, the conversions are increasing drastically as we make our voice agents more reliable and expressive. This duplex voice harness/model (interaction runtime + kernel) problem has occupied most of my attention but tbh when I first started working with voice agents, I didn't know how challenging and complex they were. Duplex conversational timing, orchestration, acoustic turn-taking, micro-interruptions, prosody, and generalized noise etc are really hard problems. There are so many 'beautifully crafted"'voice demos but seeing them in production is almost impossible. A cough would interrupt the voice agent, it leaves the caller waiting for it to speak again. During a tool call, someone would change their decision before the earlier result returned. There are also many other telephony factors so it is not just an AI or research problem. Following those cases through the system led us to separate the part that speaks from the work running in the background. We’ve developed a kernel that controls what the agent can do and checks what happened before it reports an outcome to the caller because the duplex models simply weren’t good enough yet. During this period, reading NVIDIA’s PersonaPlex and Nemotron VoiceChat work helped me understand where the models were heading; Thinking Machines’ Interaction Models approach to keeping an interaction going during longer tasks connected with questions we were already working through. My first thought was how much of what I’ve been working on will we no longer need? This is another great example of how harnesses are compressing into models. That's why always build your harness and repo that you can destroy every 6 months and build it from scratch again. It should be very fluid and modular, so you can easily bring the frontier capabilities. That's why building a duplex harness that happens to use ASR, STT, LLM, TTS models is not a good solution since duplex is commoditized today. Build a harness where the duplex model itself is a replaceable component. This is valid for any software field. After spending months with these problems, I appreciate the work behind GPT-Live more deeply. Through client delegation, anyone can connect GPT-Live to existing harnesses, where the models you choose work with your tools. Thank you to the OpenAI team. This is a big win for anyone building solutions to benefit humanity.
    @OpenAIDevsRT @TownAI: Big news: now you can strike up a conversation with your Townie, powered by @OpenAI's new GPT-Live. Tap the talk icon from anyw…
    @scottbelskykey building block for Westworld type experiences
    @mckbrandoRT @kundan2510: excited to announce the launch of gpt-live API. a few things to note: the novel architecture of the gpt-live model needed…
    @DanielleFongRT @muratcan: When I tried GPT-Live-1 today, I felt deflated as I recognized behaviors we’d spent weeks trying to make reliable. For the pa…
    @pronounced_kyleI'm gonna give mine the personality of the depressed robot from Hitchhikers Guide to the Galaxy
    @ArtificialAnlysOpenAI’s GPT-Live-1 debuts at #1 on the Artificial Analysis Speech to Speech Index at a score of 81.5 using Astra as its delegated backend model, ahead of Grok Voice Think Fast 2.0 GPT-Live-1 is @OpenAI's new full duplex Speech to Speech model that can delegate reasoning and tool use to a backend text model while continuing the conversation. Developers stream audio in and receive speech back through the API, with the backend text model configured separately. We evaluated two backend configurations: Astra at medium reasoning effort and Sol at low reasoning effort. Key takeaways: ➤ Speech to Speech Index: GPT-Live-1 (Astra, medium) achieves 81.5, ranking #1, while GPT-Live-1 (Sol, low) scores 80.1, ranking #3. Grok Voice Think Fast 2.0 High sits between them at 81.3 ➤ Speech Agent Arena: GPT-Live-1 (Sol, low) ranks #3 in preference at 1,053 Elo with 90.9% Task Success Rate, while GPT-Live-1 (Astra, medium) ranks #4 at 1,048 Elo with 87.4% task success. Gemini 3.1 Flash Live Minimal leads preference at 1,096 Elo, while Grok Voice Think Fast 2.0 High leads task success at 94.6% ➤ Tau Voice: GPT-Live-1 (Astra, medium) and GPT-Live-1 (Sol, low) take the top two spots on our agentic-performance benchmark at 67.9% and 59.3%, respectively, ahead of Grok Voice Think Fast 2.0 High at 56.5%. ➤ Big Bench Audio: GPT-Live-1 (Astra, medium) scores 90.1% and GPT-Live-1 (Sol, low) scores 89.0% on audio reasoning, behind Grok Voice Think Fast 2.0 High at 97.2% and Qwen Audio 3.0 Realtime Plus at 99.2% ➤ Speed: Average Time to First Audio on Big Bench Audio is 1.34 seconds for GPT-Live-1 (Astra, medium) and 1.24 seconds for GPT-Live-1 (Sol, low), compared with 0.70 seconds for Grok Voice Think Fast 2.0 High ➤ Cost: GPT-Live-1 (Astra, medium) costs $5.83 per hour of input audio, compared with $4.47 for GPT-Live-1 (Sol, low), including delegated backend model usage, versus $4.80 for Grok Voice Think Fast 2.0 High on our fixed Big Bench Audio pricing subset See below for more detail ⬇️
    @davidmokos_Created a GPT-Live-1 voice app template with @expo - Full-duplex audio + background calls on iOS - Choose from all supported voices - 6 Rive animations from Vercel AI Elements - Expo SDK 57 + native Expo UI controls Source below ↓
    @rohanpaul_aiAnother detail that I find architecturally interesting: it's a single Audio Encoder + LLM foundation serving both offline and streaming recognition. One base, configurable chunk sizes, and no accuracy tax on the offline side. In practice that means you're not training and maintaining two models that drift apart over time, you're tuning one system along a latency/quality axis.

    16 Sources

    @bnicholehopkinsOpenAI just shipped GPT-Live-1 in the API, and I sat down with Peter Bakkum, Member of the Technical Staff working on the realtime APIs to hear his take on how this changes the game. I have been so excited for this launch - over the past few months we've been working with Open AI in alpha testing the latest speech to speech API and I'm really impressed. Instead of chaining together separate speech-to-text, language, and text-to-speech models, GPT-Live-1 reasons across incoming and outgoing audio as one continuous interaction. That matters because VAD and turn detection have wreaked havoc on voice agents. Awkward pauses, interruptions, dropped context, background noise, and accidental responses to silence can all make an agent feel painfully robotic. GPT-Live-1 is designed around a more fluid, turnless interaction model. The result is a conversation that feels much closer to how people actually speak. It also introduces a compelling architecture for balancing natural conversation with deep intelligence. The voice model can stay fast and context-aware while delegating complex reasoning and tool calls to a backend model—including third-party models. Developers can shape tone, pace, and conversational style directly through the system prompt. Longer sessions retain context and quality more reliably. And full-duplex telephony support opens the door to real phone-based agents for customer support, restaurant reservations, scheduling, and much more. In our conversation, Peter and I cover: - What made this API particularly difficult to build and deploy - How developers can prepare to migrate from cascaded architectures - Why voice evaluations are essential - How full-duplex models change the way we build voice agents - The entirely new categories of applications this technology could unlock Voice is becoming a the interface of AI and this release brings that future meaningfully closer. Listen to the full conversation at the link below. 🎙️ Youtube: https://youtu.be/yr-Em6RL7mM Spotify: https://open.spotify.com/episode/21RZ91AMKuVbeMUrn2gknw?si=3071341125ae4ecb
    @muratcanWhen I tried GPT-Live-1 today, I felt deflated as I recognized behaviors we’d spent weeks trying to make reliable. For the past three months, I’ve been focusing on researching and building duplex voice agents, where I wrote thousands of lines across our harness and the experiments around it. Customers use parts of that work every day, the conversions are increasing drastically as we make our voice agents more reliable and expressive. This duplex voice harness/model (interaction runtime + kernel) problem has occupied most of my attention but tbh when I first started working with voice agents, I didn't know how challenging and complex they were. Duplex conversational timing, orchestration, acoustic turn-taking, micro-interruptions, prosody, and generalized noise etc are really hard problems. There are so many 'beautifully crafted"'voice demos but seeing them in production is almost impossible. A cough would interrupt the voice agent, it leaves the caller waiting for it to speak again. During a tool call, someone would change their decision before the earlier result returned. There are also many other telephony factors so it is not just an AI or research problem. Following those cases through the system led us to separate the part that speaks from the work running in the background. We’ve developed a kernel that controls what the agent can do and checks what happened before it reports an outcome to the caller because the duplex models simply weren’t good enough yet. During this period, reading NVIDIA’s PersonaPlex and Nemotron VoiceChat work helped me understand where the models were heading; Thinking Machines’ Interaction Models approach to keeping an interaction going during longer tasks connected with questions we were already working through. My first thought was how much of what I’ve been working on will we no longer need? This is another great example of how harnesses are compressing into models. That's why always build your harness and repo that you can destroy every 6 months and build it from scratch again. It should be very fluid and modular, so you can easily bring the frontier capabilities. That's why building a duplex harness that happens to use ASR, STT, LLM, TTS models is not a good solution since duplex is commoditized today. Build a harness where the duplex model itself is a replaceable component. This is valid for any software field. After spending months with these problems, I appreciate the work behind GPT-Live more deeply. Through client delegation, anyone can connect GPT-Live to existing harnesses, where the models you choose work with your tools. Thank you to the OpenAI team. This is a big win for anyone building solutions to benefit humanity.
    @OpenAIDevsRT @TownAI: Big news: now you can strike up a conversation with your Townie, powered by @OpenAI's new GPT-Live. Tap the talk icon from anyw…
    @scottbelskykey building block for Westworld type experiences
    @mckbrandoRT @kundan2510: excited to announce the launch of gpt-live API. a few things to note: the novel architecture of the gpt-live model needed…
    @DanielleFongRT @muratcan: When I tried GPT-Live-1 today, I felt deflated as I recognized behaviors we’d spent weeks trying to make reliable. For the pa…
    @pronounced_kyleI'm gonna give mine the personality of the depressed robot from Hitchhikers Guide to the Galaxy
    @ArtificialAnlysOpenAI’s GPT-Live-1 debuts at #1 on the Artificial Analysis Speech to Speech Index at a score of 81.5 using Astra as its delegated backend model, ahead of Grok Voice Think Fast 2.0 GPT-Live-1 is @OpenAI's new full duplex Speech to Speech model that can delegate reasoning and tool use to a backend text model while continuing the conversation. Developers stream audio in and receive speech back through the API, with the backend text model configured separately. We evaluated two backend configurations: Astra at medium reasoning effort and Sol at low reasoning effort. Key takeaways: ➤ Speech to Speech Index: GPT-Live-1 (Astra, medium) achieves 81.5, ranking #1, while GPT-Live-1 (Sol, low) scores 80.1, ranking #3. Grok Voice Think Fast 2.0 High sits between them at 81.3 ➤ Speech Agent Arena: GPT-Live-1 (Sol, low) ranks #3 in preference at 1,053 Elo with 90.9% Task Success Rate, while GPT-Live-1 (Astra, medium) ranks #4 at 1,048 Elo with 87.4% task success. Gemini 3.1 Flash Live Minimal leads preference at 1,096 Elo, while Grok Voice Think Fast 2.0 High leads task success at 94.6% ➤ Tau Voice: GPT-Live-1 (Astra, medium) and GPT-Live-1 (Sol, low) take the top two spots on our agentic-performance benchmark at 67.9% and 59.3%, respectively, ahead of Grok Voice Think Fast 2.0 High at 56.5%. ➤ Big Bench Audio: GPT-Live-1 (Astra, medium) scores 90.1% and GPT-Live-1 (Sol, low) scores 89.0% on audio reasoning, behind Grok Voice Think Fast 2.0 High at 97.2% and Qwen Audio 3.0 Realtime Plus at 99.2% ➤ Speed: Average Time to First Audio on Big Bench Audio is 1.34 seconds for GPT-Live-1 (Astra, medium) and 1.24 seconds for GPT-Live-1 (Sol, low), compared with 0.70 seconds for Grok Voice Think Fast 2.0 High ➤ Cost: GPT-Live-1 (Astra, medium) costs $5.83 per hour of input audio, compared with $4.47 for GPT-Live-1 (Sol, low), including delegated backend model usage, versus $4.80 for Grok Voice Think Fast 2.0 High on our fixed Big Bench Audio pricing subset See below for more detail ⬇️
    @davidmokos_Created a GPT-Live-1 voice app template with @expo - Full-duplex audio + background calls on iOS - Choose from all supported voices - 6 Rive animations from Vercel AI Elements - Expo SDK 57 + native Expo UI controls Source below ↓
    @rohanpaul_aiAnother detail that I find architecturally interesting: it's a single Audio Encoder + LLM foundation serving both offline and streaming recognition. One base, configurable chunk sizes, and no accuracy tax on the offline side. In practice that means you're not training and maintaining two models that drift apart over time, you're tuning one system along a latency/quality axis.