Nuance Labs, a Seattle-based startup building AI models for face-to-face conversations, has closed a $50 million early-stage round. The Series A was led by Lightspeed Venture Partners, which had already backed the company's seed round, with Accel and South Park Commons returning and Nvidia Corp. and Define Ventures joining as new investors. The money goes toward a model that lets avatars respond while a person is still speaking, rather than after a pause.
What the company is building
The round is the second check Lightspeed has written for Nuance, and the investor list now includes a chipmaker whose hardware sits under most AI workloads. Co-founder and Chief Executive Fangchang Ma, formerly an AI researcher at Apple Inc., said the company is building a human foundation model rather than another layer on top of existing chatbots. The team has published a demo of the model, which Ma describes as still a work in progress. No product has shipped yet: Nuance hopes to release its first research preview to the public later this year, and the new funds will accelerate development and pay for more researchers.
Existing voicebots and avatars are assembled from separate parts, Ma told Business Insider: a voice-to-text model, a large language model that writes the answer, and a text-to-voice model that turns it back into audio. Each handoff adds delay, and an avatar needs a fourth step to animate a face that matches the words. Nuance instead runs one full-duplex system that takes in audio and video of the user and streams audio and video back at the same time. The model reads words, gaze, gestures, tone and timing, and answers with facial and vocal expressions in real time, learning from how people behave during conversations.
The awkwardness Ma points to is a known limit of the current stack. Because the pieces work in sequence, an avatar that is listening shows no reaction, which breaks the sense of talking to a person. Nuance's bet is that perception and response belong in one model, not in a pipeline. The company names sales and customer service, coaching, professional training and education as target uses, where expressions can carry part of the result, for example a video call to practice a language or rehearse an interview.
What this means for business
For companies already running AI agents in support or sales, the practical change is the shape of the interaction, not the model's benchmark scores. A system that reacts while the customer is still talking can shorten a call and reduce the interruptions that make voice bots feel mechanical. Small teams that buy ready-made avatar tools will notice this first in customer-facing calls; larger organizations will weigh it against the cost of rebuilding existing pipelines, where a text-based agent already handles most requests without any video at all.
What the news does not mean is that natural avatars are available now. Nuance has no product, only a demo and a stated plan for a research preview later this year, so any deployment decision rests on that release and on how the model performs outside controlled conditions. Buyers should ask how the system handles accents, poor audio, and privacy of video streams, and whether it can be integrated with existing contact-center software. The $50 million funds research, not a shipping roadmap with dates.
The marker to watch is the research preview promised for later this year: if it opens to outside developers with an API and holds up in real conversations, the full-duplex approach becomes a layer others build on, as Lightspeed's Nnamdi Iregbulem expects. If the preview slips or stays closed, avatars remain a pipeline of separate models, and the delay Ma describes stays part of the product.
