Anthropic released a major update to Claude Voice on July 23, expanding model choice beyond Haiku to Opus and Sonnet, enabling multi-turn voice reasoning and complex task execution. Voice mode now defaults to the last model a user selected in text chat and allows mid-conversation switching via model picker. Users can ask Claude to draft emails, reschedule calendar meetings, update Notion documents, or search Gmail via voice, with tool access parity to text chat (Free tier: Haiku + 1 connected app; paid: Opus/Sonnet + all connected services including Gmail, Google Calendar, Google Docs, Slack, Canva, Notion). The company added 18-language multilingual support, though manual language selection is still required.
Architecturally, Claude Voice uses a turn-based model: Claude listens, pauses to think, then speaks—unlike OpenAI's bidirectional GPT-Live (available to all ChatGPT users since July 8), which processes speech and generates output simultaneously for a more conversational cadence. Anthropic intentionally chose stronger reasoning models layered onto conventional voice infrastructure (reportedly ElevenLabs for text-to-speech) over building a speech-native system. The company stated this release is "focused on intelligence and tool access," signaling a deliberate tradeoff: superior task reasoning and connected-app integration at the cost of natural interruption handling and latency.
For enterprise adopters, Anthropic's bet is that tool-connected reasoning matters more than conversational fluidity. Anthropic derives ~80% of revenue from 300K+ enterprise customers, so voice-to-Gmail/Slack integration directly impacts productivity workflows. OpenAI's bid (full-duplex, hands-free agents that spin up background tasks) targets different use cases: brainstorming, delegation without back-and-forth. Neither company has yet shipped an assistant pairing speech-native naturalness with broad tool access. Monitor which approach resonates with users as deployments scale; the answer will shape whether voice AI prioritizes inference model quality or audio infrastructure. Anthropic's "more to share later this year" suggests the company plans to address the voice-naturalness gap, though timeline and approach remain undisclosed.