Many readers have likely experienced the frustration of trying to chime in with an acknowledgment while talking to a voice assistant, only to be made to wait until the AI finished its response. In human conversation, exchanges flow naturally with interjections like "uh-huh" and "right, right," but conventional voice assistants have mostly been designed to stay silent until the other party finishes speaking before responding—a one-way street. On July 8, 2026, OpenAI announced "GPT-Live," a full-duplex voice model that breaks down this one-directional design, and made voice conversations powered by GPT-5.5 (OpenAI's large language model) available for free to free-tier ChatGPT users as well. For standalone voice AI companies that have relied on pricing as a competitive weapon, this free rollout is a move that could shake the very foundation of their business.
How Full-Duplex Voice Conversation Works: The Mechanics of GPT-Live
OpenAI announced GPT-Live on July 8, 2026, and began its global rollout for ChatGPT users. GPT-Live adopts a full-duplex architecture that allows the model to listen simultaneously even while the user is speaking, enabling it to return acknowledgments or interject in real time. Most conventional voice assistants have used a half-duplex approach that alternates between "listening" and "speaking," with the model waiting until the user's speech pauses. OpenAI explains that GPT-Live eliminates this waiting time, allowing the conversation to continue while responding to backchannel cues like "mhmm" and "yeah."
The underlying models are divided by use case. Free users are assigned the lightweight GPT-Live-1 mini by default, while paid users on Go, Plus, and Pro plans receive the higher-tier GPT-Live-1. Both of these run on top of GPT-5.5 Instant. For situations requiring deeper reasoning, GPT-Live-1 Medium and GPT-Live-1 High are assigned, using GPT-5.5 Thinking with reasoning effort set to medium or high, respectively. The supported platforms are iOS, Android, and ChatGPT.com.
Free users are assigned the lower-tier GPT-Live-1 mini model, which likely differs in response quality from the paid version, GPT-Live-1. Even so, the full-duplex experience itself—listening and responding with acknowledgments while the user is still speaking—is offered to the free tier under the same design philosophy. The rarity of the "natural, human-like conversation" experience that standalone voice AI companies have marketed as a premium feature is already being eroded, starting from the entry-level model.
When a query requiring web search or complex reasoning comes in, GPT-Live is designed to hand off processing to the latest frontier model behind the scenes and then weave the result back into the flow of conversation. OpenAI claims GPT-Live significantly outperforms the previous-generation Advanced Voice Mode on GPQA (a benchmark for specialized reasoning), BrowseComp (a benchmark for web search), and τ³-Voice Telecom (a benchmark simulating customer support scenarios), though specific scores and margins of improvement are not listed on the official page. While the model supports multiple languages, OpenAI itself acknowledges that non-native-sounding pronunciation remains in some languages, and TechCrunch noted that a Hindi translation shown in the demo had unnatural pronunciation and an overly literary tone.
The Scale Behind 150 Million Weekly Voice Feature Users
OpenAI's announcement materials state clearly that "more than 150 million people talk to ChatGPT using Voice or Dictation every week." This figure of 150 million refers specifically to users of voice-related features, not the total weekly active users of ChatGPT as a whole. Various reports put ChatGPT's overall weekly active users at over 800 million. If that figure is used as a benchmark, the calculation suggests that fewer than one in five overall users are using voice features.
The free rollout of GPT-Live is a measure that enhances the experience for the 150 million people who already use voice, and it does not directly change anything for the majority of ChatGPT users. Even so, this scale alone is a level that standalone, voice-focused paid apps would struggle to reach even combined. The order of magnitude of the customer base that standalone voice AI companies are aiming to capture has, with this free rollout, been decisively surpassed.
Why Video and Screen Sharing Were Left Out, and the Gap with Gemini Live
OpenAI's official announcement includes a single line stating: "At launch, GPT‑Live will not support voice with video or screen sharing in ChatGPT." Compared to the detailed explanations of full-duplex conversation quality and language support, this line is remarkably brief, mentioned only in passing amid descriptions of other features. The lightness with which this limitation is treated suggests that OpenAI does not consider this constraint a major shortcoming.
Sustaining a full-duplex conversation requires continuously processing the user's speech, acknowledgments, and silences on a millisecond-by-millisecond basis. Attempting to simultaneously analyze video information on top of that would likely come at the cost of response latency and naturalness. Given that GPT-Live's biggest selling point is natural conversation, the lack of video support reflects a design decision that prioritizes voice latency and fluidity.
Google, meanwhile, has already adopted a native multimodal approach with Gemini 3.1 Flash Live, launched on March 26, 2026, in which a single model directly processes both audio and video and responds via voice. Even if GPT-Live leads in conversational smoothness, Gemini Live has the upper hand for use cases where users want to talk while showing something on screen. OpenAI's bet on focusing exclusively on voice does not reach users seeking a conversational experience that includes video.
What the Shift from Realtime to GPT-Live Reveals
The name of OpenAI's voice AI has changed at least four times over the past year and a half. The "Realtime API," which debuted in 2024, was revamped into "gpt-realtime" in August 2025, reportedly improving Big Bench Audio accuracy from 65.6% under the previous model to 82.8%. It subsequently transitioned to the "gpt-realtime-2" lineage, and on July 6, 2026, gpt-realtime-2.1 and gpt-realtime-2.1-mini had just been released. Only two days later, on July 8, a new consumer-facing brand, GPT-Live, was introduced.
Behind the frequent renaming lies not only the rapid pace of performance improvements but also an ongoing proliferation of product lines. On the consumer side, there are four tiers in the GPT-Live-1 lineup (mini, standard, Medium, and High), while on the developer side, there are two variants in the gpt-realtime-2.1 lineup (standard and mini)—together amounting to six voice models being supported simultaneously. In fact, the GPT-Live API (Application Programming Interface) for developers and enterprises has not yet launched; at this stage, OpenAI is only accepting sign-ups for notifications. The structure—rolling out the free consumer offering worldwide on the very same day while deferring the developer API—indicates that OpenAI is prioritizing consumer adoption. Developers who have already built voice applications on the existing Realtime API lineup will need to continue using the gpt-realtime-2.1 lineup for the time being, until the GPT-Live API becomes available.
Pricing information for the existing Realtime API lineup cannot be confirmed on the official pricing page, but according to third-party aggregator sites, GPT-Realtime-2's voice input is reportedly priced at $32 per million tokens (with cached input at $0.40), and voice output at $64 per million tokens. Converted at the exchange rate as of July 8, 2026 (1 USD = 162.43 JPY), this comes to roughly ¥5,200 for input and roughly ¥10,400 for output. If this pricing is accurate, the free rollout of GPT-Live can be interpreted as a strategy to first lock in a free consumer experience on a separate track from the paid developer API.
The Line That Will Divide Winners and Losers in the Voice AI Market
The biggest beneficiaries of GPT-Live's free rollout are free users who can now access full-duplex voice powered by GPT-5.5 at no additional cost, and OpenAI itself, which locks in that user base. If the same experience were sought via the developer API, GPT-Realtime-2's voice input would cost roughly ¥5,200 per million tokens and voice output roughly ¥10,400 per million tokens. What free users have gained in hand is precisely the functionality that would otherwise have required paying that per-unit price. TechCrunch has pointed out that multiple companies—including Apple, Amazon, and Sesame—are pursuing the same direction: natural, conversational voice UI.
Among them, standalone voice AI startup Sesame raised $250 million (roughly ¥40.6 billion) in a Series B round in October 2025, bringing its total funding to $307.6 million (roughly ¥50 billion), with a valuation reportedly exceeding $1 billion. Its demo drew over one million participants, with cumulative conversation time exceeding five million minutes—but this scale reflects usage of a beta or demo version, and conversion to paid contracts was still ahead. Now that comparable full-duplex conversation is available for free through GPT-Live, Sesame must justify the value of a product built with massive funding on the same footing as an experience that is offered for free to ChatGPT's free-tier users.
Others stand to lose as well. Google's Gemini Live also loses one point of differentiation now that ChatGPT's free tier has caught up in voice quality. Only 104 days separate the launch of Gemini 3.1 Flash Live (March 26, 2026) from GPT-Live's free rollout (July 8, 2026), leaving Google limited time to establish its strength in simultaneous voice-and-video processing among consumers. That said, Gemini Live retains an advantage over GPT-Live in being able to process voice and video simultaneously, and it remains outside the scope of OpenAI's free-tier strategy for conversational experiences that involve video or screen sharing.
The decision to demote voice conversation from a differentiating feature to a free, standard one has the effect of undermining the raison d'être of standalone voice AI companies. At the same time, the deliberate exclusion of video and screen sharing means OpenAI has drawn its own boundaries around this voice-focused bet. The timing of the developer API's launch and when video support is eventually implemented will serve as the next indicators for judging whether this bet pays off.
