Enterprise voice AI has full-duplex models that can speak while listening, yet it still lacks a ChatGPT moment of natural, trusted conversation. That assessment came from PolyAI CTO Shawn Wen and Otter CMO Alex Gay at the HumanX conference last month. Investors keep pouring billions into model makers, customer service providers, meeting notetakers and dictation tools, but weekly releases that promise human-like dialogue still fall short in practice.
Why full-duplex alone does not fix conversation
PolyAI operates an enterprise voice platform for customer service, while Otter develops a meeting notetaker with transcription and summaries. Wen said the industry has reached the milestone of full-duplex models, and the next challenge is very fast reasoning so models fetch answers quickly. Gay pointed to speaker identification, intent capture and combining transcripts with organizational knowledge as the layer that enables automation. Otter is also working on digital twins that could represent people in meetings.
The mechanics of the gap start with recognition and speed. Wen argued that ASR models often miss important keywords, which prevents the system from capturing full context. Gay agreed that transcription was never the end point for Otter, only the foundation for productivity gains and follow-up actions. If accuracy is insufficient, summaries and automated actions become flawed, and one wrong action destroys trust in the platform. Fast retrieval plus precise transcription are therefore preconditions for delegation.
The discussion reflects a crowded funding cycle around voice as the next interface. Capital flows into model developers, enterprise contact-center suppliers, note-taking products and dictation startups, with new human-sounding releases appearing almost every week. At the same time, assistants still misunderstand users and notetakers produce wrong transcripts or summaries. The contrast between marketing claims and operational reliability explains why executives describe progress as incremental rather than transformative.
What this means for business use of voice
For companies using voice agents in service and sales, the practical test is the first two or three turns of a call. Wen said agents should not sound robotic and must give callers confidence that problems will be solved. Once voice quality is good enough to sustain that opening, customers build confidence over time and may stop demanding a human handoff. That changes staffing and routing only when containment holds without escalation, which favors large contact centers first and smaller teams after proven templates emerge.
The second consequence concerns meetings, records and disclosure. Gay said digital twins need the same emotive expression as human discussion, otherwise debate and strategic conversation collapse into a question-and-answer chatbot. Otter wants to notify all participants in chat that recording continues even when its bot is absent, while Wen stressed that enterprise callers must know they speak with AI. Buyers should therefore verify keyword accuracy, multilingual coverage, speaker attribution and explicit notification controls before allowing autonomous actions.
The marker to watch is whether vendors demonstrate fast reasoning combined with measurable transcription accuracy in live deployments. Further signals include stable containment after three turns in service calls and declared AI identity with recording notices in meetings. If those appear together across PolyAI-style platforms and Otter-style assistants, voice will move from demonstration to dependable business infrastructure.
