On July 23, 2026, Anthropic made Claude Opus and Claude Sonnet available in Claude's voice mode for paid plans. Voice mode on Free continues to be limited to Claude Haiku, but on paid plans, depending on the subscription, the mode now carries over the model family last selected in text conversation and allows switching mid-conversation. Voice can also now reach connected tools like Gmail, Slack, and Google Calendar, and the feature supports 11 languages and regional variants, including Japanese. As voice moves closer to matching the model selection and tool use available in the text version, it also brings new constraints: turn-based conversation, plan-specific usage limits, permission confirmations, and stored data.

AD

Removing the Haiku lock-in, choosing models mid-conversation

The reason the old voice mode used Haiku was speed. According to Anthropic, users began using voice not just for quick questions but to think through longer work-related problems out loud, and this created situations where conversations were fast but lacked depth. To address this, Anthropic added Opus, which handles complex reasoning, and Sonnet, used for writing, analysis, and multi-step tasks, as options.

What users select are family names—"Opus," "Sonnet," or "Haiku"—rather than specific generations. When starting voice mode, the system picks up the family used in the most recent text conversation and automatically moves to the latest generation within that family. Anthropic explains that to keep conversations smooth, it uses a fast version of the selected model. If a consultation that started with Haiku becomes more complex, users can switch to Opus while preserving the conversation.

However, the model lineup for text and voice doesn't match exactly. In Free text conversations, users can access Haiku and Sonnet, but voice mode is limited to Haiku. On paid plans, depending on the subscription, this extends to Sonnet and Opus, but Fable—the model designed for long-duration autonomous work—is not supported in voice mode.

Model changes also affect usage time. Voice conversations consume the regular usage limit, and Anthropic's model selection guide explains that Haiku consumes the least of this limit, Sonnet a moderate amount, and Opus the most. While users gain access to deep reasoning via voice, the amount of conversation they can sustain within the same usage allowance may decrease. Switching models is also, in effect, a choice between performance and remaining time.

Continuing to think aloud while keeping the turn-based structure

Claude's voice mode is not a system where the user and model speak simultaneously in real time. Claude listens to what's said, pauses to think, and then responds—a turn-based structure. Anthropic has kept this pause intact while using Opus and Sonnet to increase follow-up questions and deeper exploration of points, pushing forward use cases like finding gaps in a plan or comparing multiple hypotheses.

There are two modes of operation. In the default hands-free mode, the system detects natural pauses and begins responding; if the user starts speaking, Claude stops talking and listens again. In noisy environments, users can switch to push-to-talk, which only transmits audio while a button is held down. Within the same thread, users can switch back to text to enter a URL or code, then return to voice, and the prior conversation carries over.

The feature is available on iOS and Android mobile apps, Claude Desktop, and the web, as a beta for all chat users. Anthropic itself notes that it works best on smartphones. Speech recognition accuracy, time to first voice response, and interruption success rate have not been disclosed. It cannot yet be confirmed that "switching to Opus improved the naturalness of conversation itself"—what has been confirmed with this update is the model family handling the conversation and the scope of tasks it can perform.

AD

A voice interface that reads email and moves appointments

The new voice mode can call on connected Gmail, Google Calendar, Google Docs, Slack, and Canva during conversation. In examples Anthropic gives, users can ask by voice to summarize the day's emails and draft replies to important ones, push back a meeting by 30 minutes when running late, or turn a conversation about a client proposal into a one-page Canva document.

Connecting to external data itself is not new this time. When Anthropic released the mobile voice mode on May 27, 2025, it also gave examples like summarizing Calendar and searching Docs. What's different this time is that tasks can now be handed off from Haiku to higher-tier models, the range of connected services has expanded, and actions that modify appointments or documents after a search are now built into the same voice conversation.

Being able to ask for something by voice is not the same as being able to execute it unconditionally. Claude asks for permission before using connected tools and does not access data beyond the permissions the user holds in the original service. In Google Calendar, it can create, update, and delete events, but for Gmail, it's limited to search and drafting—sending emails still requires the user to do so within Gmail itself.

Organizations also have their own stopping points. Team and Enterprise owners can set per-action controls—"always allow," "require approval," or "block"—such as permitting Google Drive reads while blocking document creation or editing. Free users can use one connected tool in voice mode, while paid plans can use all connected tools. As voice-based operations increase, what matters most during setup is not the model name but the permission table specifying which actions are allowed for whom.

11 languages including Japanese, with manual switching

Voice settings now support 11 languages, including English, Japanese, and Korean. Hindi and Indonesian have also been added. Among European languages, French, German, and Italian are supported, while Portuguese is the Brazilian variant. Spanish has been split into Latin American and European Spanish versions. Language support is available across all plans.

Support for languages other than English is in beta, and Claude does not automatically detect the spoken language. Users need to select a language in voice settings or explicitly instruct a switch during conversation. Even if the app's display language is set to Japanese, this setting is not carried over to voice settings. What matters practically is not the number of supported languages, but whether context and proper nouns are preserved when switching languages mid-topic.

AD

The management boundary lingering in voice recordings and connected data

Text transcripts of voice conversations are saved to chat history just like regular chats. Data retrieved from connected tools is also stored on Anthropic's servers alongside the related chat, and can be deleted by deleting the chat. Google Workspace connector data itself is not used for model training, but if a user manually pastes content and has training-use permission enabled on a consumer account, it may be treated differently.

There are differences depending on the specific feature involving voice input. Anthropic explicitly states that for the dictation feature in its commercial products, voice recordings are deleted once transcription is complete, and voice is not used for model training. Voice mode's current FAQ, on the other hand, only explains the retention of text transcripts—whether voice recordings themselves are retained, and if so, when they are deleted, is not stated on the same page. The dictation feature's handling policy cannot be assumed to apply to voice mode.

Output voice is also not designed for free reproduction. Users choose from a limited set of preset voices, and there is no feature to clone or mimic a specific person's voice. Additionally, while dictation is available in Claude Cowork and Claude Code, voice mode is not available in either. Voice mode also cannot access Projects or Skills placed in Cowork.

With this update, Claude has gained a pathway for thinking through difficult problems by voice and transferring conclusions to connected tools. The next evaluation won't be decided by general model performance charts. Whether recognition accuracy during extended Japanese speech, response time after switching to Opus, success rates for rescheduling appointments or creating documents, and voice recording retention conditions are disclosed—and whether they hold up in real-world conversations—will determine how widely this spreads.