Skip to content

Product1 publisher2 min readPublished

Nuance Labs raises $50M to replace the avatar's four-model pipeline with one

Fangchang Ma says the lag and the dead-eyed listening in voice products come from the handoffs between four models. Nuance's evidence is a demo it calls unfinished and a research preview promised for later this year.

The Product Desk · Product desk

Photograph accompanying Nuance Labs raises $50M to replace the avatar's four-model pipeline with one
Photo: scour.ing

What happened

  • Nuance Labs, a Seattle startup, said it closed a $50 million early-stage Series A led by Lightspeed, with Nvidia and Define Ventures new to the cap table and Accel and South Park Commons returning.
  • Chief executive Fangchang Ma, formerly an AI researcher at Apple, told Business Insider that today's avatars chain a voice-to-text model, a language model, a text-to-voice model and a face animator.
  • Nuance's pitch is a single full-duplex model that takes audio and video in and streams audio and video back, reacting to a user's words, gaze, gestures, tone and timing while they are still talking.
  • The company has not launched a product and hopes to put a first research preview in public hands later this year, with the new money going to development and more researchers.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • decision Anyone with a conversational feature due this quarter now has to price the option of waiting for a single-model vendor whose earliest public artefact is a research preview with no date beyond later this year.
  • constraint A team cannot justify that wait on measurement, only on Ma's explanation of why the handoffs hurt.
  • capability If reaction during the user's own turn works as Ma describes, interview rehearsal and coaching become buildable; a refund-status bot only needs the answer.
  • cost Moving to one model means retiring four components at once, along with the evals and vendor arrangements a team built around each of them.

Fangchang Ma, Nuance Labs' co-founder and chief executive, told Business Insider that the face that stops moving three words into a user's sentence comes from how the products are put together, not from a fault in any one part [5][8]. "If you use existing AI avatars, when they're listening, they're not reacting, or they're just doing random things," Ma said [9].

A typical avatar runs four stages. Speech to text, then a language model, then text to speech, then a face animated to match the audio [6]. Ma said replies come back late because the system processes everything step by step [7]. Four stages means three handoffs, and Nuance's answer deletes all three: one model, audiovisual in and audiovisual out, perceiving the user's stream while streaming a response back [10][19]. "We decided to build this on one system, on one model," Ma said. "There's audio/video in and audio/video out." [12]

What exists today is a demo the company describes as a work in progress [13]. The account includes no latency measurement and no benchmark against the assembled pipeline [20]. So the case for one model over four rests on Ma's explanation and on that demo. Nnamdi Iregbulem of Lightspeed, which also backed the seed round [2], said of the founders that "they've turned that into a single model that follows you in real time" [17].

Ma is aiming at sales and customer service, coaching, professional training and education, including practising an interview or learning a language over a video call [15]. Those split on one question: where the value arrives. A customer asking when a refund clears wants the answer, and a beat of dead air before it costs them almost nothing. Someone rehearsing a job interview is buying the listening half, the nod at the right moment and the flicker when an answer rambles. Ma said the model can demonstrate understanding in the moment, while the user is still speaking [11].

That question decides whether waiting is worth it. Where the value lands in the reply, the assembled pipeline can carry the product, and shipping it now is the honest call. Where the value lands during the user's turn, the pipeline you wire together cannot deliver it, and the vendor making the claim has a preview coming and no product [14]. Either way, the useful figure to have before that preview arrives is task completion and repeat use on what you run now, because that preview can only be measured against your own baseline.

What to watch

  • Whether the research preview lands this year, and whether it ships with a latency figure anyone outside the company can reproduce.
  • Whether the first release is a developer API with a price or a hosted avatar product, and which of Ma's named use cases it covers.
  • What Nvidia's participation brings beyond the check, such as hardware access or distribution for the model.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories