On July 23, 2026, OpenAI made Voice available in Work and Codex within the ChatGPT desktop app. For eligible accounts, users can now speak to start tasks, interrupt work in progress, and coordinate multiple agents. Voice has expanded from a conversational feature that reads out responses into a control surface for driving agent work. That said, the tools and operating permissions Voice can use remain bounded by whichever Work or Codex context is selected, and usage time and task execution are tracked in separate allotments.
From Chat to Work/Codex: Voice's Role Shifts
This addition sits at the junction of three updates rolled out over July. On July 8, OpenAI introduced GPT-Live-1 to consumer-facing ChatGPT Voice. Paid plans use GPT-Live-1, while Free uses GPT-Live-1 mini, with the model processing listening and speaking simultaneously. Users can now interrupt mid-response, and while listening to voice, they can also follow along with text streamed in the same chat.
The next day, July 9, brought the announcement of ChatGPT Work—designed for long, multi-step jobs—along with a new desktop app. Work handles research and analysis, producing finished artifacts like documents, spreadsheets, and presentations. Codex, meanwhile, takes on code writing, debugging, testing, and repository work. Chat, Work, and Codex kept their distinct roles while landing in the same app on macOS and Windows.
On July 16, a screen for switching between Chat and Work went live, along with a history that consolidates conversations from both. Projects also entered the app, allowing Work sessions in the cloud to be picked up across devices. The July 23 update connects Voice to that execution layer. GPT-Live-1 handles natural conversation, while Work and Codex carry out the actual work. More significant than simply swapping voice models is the shift toward handing off work from conversation to agents—and then being able to intervene by voice afterward.
From Starting Tasks to Coordinating Multiple Agents
Voice in Work and Codex continues working after the initial spoken instruction. Users can start a task, reprioritize it, pause it, or redirect it. Work continues in the background, and Voice reports blockers or completion either audibly or on screen. It's a design that lets you delegate long-running jobs and supervise them without returning your hands to the keyboard.
This also covers operations spanning multiple agents. Voice moves across active conversations and Projects, helping coordinate which task should take priority. Given the right project context and connected tools, it can even resume work involving documents or calendars—along with contacts and communication methods. For example, you might check on a Work session preparing for a meeting alongside a Codex session advancing related implementation, then pause one while adding conditions to the other—the kind of oversight this is built for.
While operating, you can see on screen whether ChatGPT is listening, and you can mute or stop the microphone. Live's responses are displayed as text in parallel with being spoken, and remain viewable in history afterward. However, Voice transcripts are not verbatim records. When speech overlaps or background noise intrudes, the transcript can diverge from what was actually said—so approved terms and task conditions should be verified against the on-screen record.
Before Screen Sharing, There's Permission Design
OpenAI describes Voice in Work and Codex as being able to operate your computer. This doesn't mean Voice has gained some new universal permission. Voice uses the tools and permissions already available within whichever Work or Codex context is currently selected. Voice is the entry point for operation; it's the agent-side settings that determine the scope of execution.
Getting started requires microphone access. When handling context on your computer, macOS or Windows Screen & Audio Recording and Accessibility permissions may also be needed. Whether Work can access local folders or desktop apps depends on your plan and workspace settings. Local files and outputs stay on that computer unless the user explicitly moves or shares them.
Here it's important to distinguish between Live's "screen sharing" and Work/Codex's "computer operation." GPT-Live-1's Live initially didn't support video or screen sharing—those are features for eligible iOS/Android users via Advanced Voice. Desktop Voice, by contrast, invokes Computer Use and connected tools already permitted to Work/Codex, by voice. This is distinct from continuously streaming desktop video into a Live voice conversation.
There are boundaries on availability too. Voice in Work and Codex is a macOS and Windows desktop feature and cannot be launched standalone from web or mobile. A path exists for remotely accessing desktop work from a paired iOS device, but that's not the same as selecting Codex itself on mobile. Also, only one Voice conversation can be open at a time.
About 6 Credits per Minute for Business/Flexible-Billing Enterprise
For Business and flexible-billing Enterprise plans, time spent connected via voice and work executed by agents are metered separately. ChatGPT Business workspaces include one hour of Voice in Chat, with additional usage costing 5 credits per minute. Voice in Work and Codex consumes about 6 credits per minute. Furthermore, tasks delegated via Voice are deducted at standard rates from the shared agent usage pool for Work and Codex.
This two-tier structure clarifies the difference from ordinary Chat Voice. Time spent in ongoing Chat conversations falls under regular voice usage allotments and doesn't consume Codex usage. When using Voice within Work/Codex, connection time is metered separately, and any actual work initiated from there draws from the shared agent usage pool. For workflows that keep voice operation open for extended periods, you'll need to budget for both conversation time and delegated work.
For flexible-billing Enterprise as well, Voice in Chat runs 5 credits per minute, and Voice in Work and Codex runs about 6 credits per minute. Traditional Enterprise plans include roughly 45 minutes of Voice in Work and Codex per 5-hour block. It hasn't been disclosed whether these same figures apply across all plans, including individual/consumer plans. Availability varies by plan and region, and is also affected by workspace settings and app version.
For Enterprise, Edu, and Healthcare, during the first two weeks after rollout begins, admins must enable both "Advanced voice capabilities" and "Early Model Access" before Live can be selected. Existing Advanced Voice conversations won't automatically switch over to Live either. For those rolling this out, the work of defining role-based access controls and usage caps comes before choosing a model.
Can Voice-Driven Instructions Be Audited?
Operating agents by voice also changes how you read the operation history. Voice clips from Live and Advanced Voice are retained for 30 days, alongside the transcripts shown in chat history. Deleting a conversation generally deletes associated clips within 30 days as well, though archiving does not count as deletion. Given that spoken content and text transcripts may not match, it's more reliable to keep important conditions as text within the task itself.
Handling for training use differs between voice/video clips and content like transcripts. Voice and video clips are not used for model training unless the user explicitly opts to share them. Business, Enterprise, and Edu workspaces cannot share Voice's audio/video clips for training purposes in the first place. Transcripts and attachments, on the other hand, follow plan and data settings—so administrators should check voice retention settings and content-usage settings separately.
Voice in Work and Codex significantly reduces the friction of handing work off to agents. At the same time, it also raises the speed at which ambiguous verbal instructions can propagate across multiple tasks. Even after completing the required enabling steps in the first two weeks, practical considerations remain. Only once you've determined which tools to permit for what, how to manage voice time and task volume, and which text should preserve approval conditions, does voice become a safe control surface.
