Nuance Labs raises $50M to teach AI avatars how humans actually talk
Lightspeed led the round for the ex-Apple researchers, whose full-duplex audiovisual model is due for a public preview later in 2026.
By Ryan Merket · Published
Primary source: X - Fangchang Ma
Why it matters
Nuance Labs is spending foundation-model money on the missing mechanics of face-to-face AI: timing, gaze, tone and interruption. Its 2026 preview will show whether that requires a new model or merely better product engineering.

Nuance Labs co-founder and CEO Fangchang Ma (@fangchangma) announced a $50 million Series A on September 14th to develop an audiovisual foundation model that can interpret and generate speech, facial expression and body language in real time.
Lightspeed Venture Partners led the financing, with participation from Accel, South Park Commons, NVIDIA and Define Ventures. Accel, Lightspeed and South Park Commons previously backed Nuance Labs' $10 million seed round in 2025, bringing its announced funding to $60 million. NVIDIA and Define Ventures are new investors, according to Nuance Labs.
The financing had been underway well before Monday's public announcement. A March regulatory filing reported that Nuance Labs had sold about $46.7 million of securities toward an offering capped at roughly $60.8 million. The filing listed March 12th as the date of the first sale and named Lightspeed partner Nnamdi Iregbulem as a director.
Ma founded the Seattle research lab with Edward Zhang and Karren Yang after the three worked as researchers at Apple. Ma earned a PhD in robotics and machine learning from MIT, while Zhang completed a PhD in computer graphics at the University of Washington. Ma and Zhang contributed to the Digital Persona system for Apple's Vision Pro before entering South Park Commons' Founder Fellowship in 2024.
Yang, Nuance Labs' chief scientist, also earned a PhD from MIT and worked on audiovisual generation at Apple. Nuance Labs now lists more than 20 researchers, engineers and operations staff on its website, up from the four-person operation described when the seed round was announced in September 2025.
One model instead of an avatar assembly line
Nuance Labs is pursuing a specific technical response to the awkward pauses, interruptions and fixed expressions common in AI avatars. Most avatar systems pass a conversation through several components: speech recognition converts audio to text, a language model writes an answer, a speech model produces a voice and an animation system moves the face.
Each handoff can discard information and add latency. Tone, hesitation, gaze and gestures may never reach the language model, while the avatar often remains static until the pipeline finishes producing a response.
Nuance Labs says its model processes incoming audio and video while simultaneously generating an audiovisual response. This full-duplex architecture is designed to let an avatar react with verbal and nonverbal cues while a person is still talking, rather than waiting for a completed transcript and taking turns.
"The potential for AI and avatars to enhance our lives will never be achieved while they stare at you blankly or keep interrupting you," Ma said in the funding announcement.
The approach builds directly on the founders' work reconstructing people for telepresence at Apple. Nuance Labs is betting that higher visual fidelity alone cannot make a digital person convincing. The underlying model also has to follow timing, expression and conversational feedback as they unfold.
Nuance Labs has identified sales, customer service, coaching, professional training and education as possible applications. The breadth of that list reflects the platform ambition behind the round: Nuance Labs wants to supply an interaction layer that other products can build on, alongside any applications it develops itself.
The public test comes later
Nuance Labs has yet to put the model into broad public use. Ma directed prospective collaborators to a waitlist for an early research demo, and Nuance Labs plans to open a public research preview later in 2026.
That preview will provide the first meaningful test of the central technical claims. Nuance Labs will need to demonstrate that one audiovisual model can respond quickly enough for live conversation while preserving identity, expression and coherent speech. It will also have to show that reading facial and vocal cues produces a useful interaction, rather than another polished avatar demo.
The Series A is earmarked for model development, research hiring and the preview. Nuance Labs is recruiting across pretraining infrastructure, reinforcement-learning research, model optimization, inference, data infrastructure and human evaluation, according to its careers page.
The hiring plan shows where the capital is going. Nuance Labs is building the training, serving and evaluation stack required to operate its own specialized foundation model, an expensive route compared with placing an animated face over an existing voice API. Lightspeed's return as lead investor amounts to a larger bet that Ma, Zhang and Yang can turn their Apple research into a model architecture that survives contact with real conversations.