Skip to content

Product1 publisher3 min readPublished

Reporters asking Synthesia the basics can now talk to an avatar of its corporate affairs head

Synthesia, valued at $4 billion, has an interactive avatar of its corporate affairs head answering common press questions about the company. Comms leads thinking of doing the same need a rule for what a synthetic stand-in may say to outsiders.

The Product Desk · Product desk

Photograph accompanying Reporters asking Synthesia the basics can now talk to an avatar of its corporate affairs head
Photo: techcrunch.com

What happened

  • Alexandru Voica, Synthesia's head of corporate affairs, sent a TechCrunch reporter a link to an interactive avatar of himself, trained to answer common press questions about the company.
  • According to TechCrunch, the avatars Synthesia then built of the reporter were the first it had made for anyone other than Voica.
  • Her interactive avatar answers questions about one of her stories only: why venture-backed startups commit more fraud than startups without VC backing.
  • Synthesia sells enterprises interactive training video with avatars, and recently launched Roleplay Sessions, which scores employees on practice sales pitches.
  • Synthesia reached a $4 billion valuation earlier this year, and last year it said it had passed $100 million in annual recurring revenue.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • decision Comms and enablement leads now have a live example to point to, so they need a written rule on whether an executive's avatar may answer outsiders at all.
  • exposure The executive who lends a face and voice is tied to every answer the avatar gives, including answers to reporters that executive never spoke with.
  • precedent With the vendor's own corporate affairs head using the product on the press, customer comms teams have an easier case for proposing the same for their executives.

The reporter showed her script-reading avatar, reciting a few lines she typed about fall arriving in New York, to some friends outside tech. They "found it both interesting and creepy," she wrote [15]. Before she met her own avatar she was indifferent to them. Now she has "no qualms" presenting it to readers [16].

Synthesia has three product lines. There is a video platform where avatars read scripts, a Sessions platform for surveys and roleplay, and an API for customers who want to build their own interactive avatars or other products [5]. The press avatar is the vendor's own comms head putting his face on his own job [1]. The article does not say how many reporters have used Voica's avatar or whether Synthesia sells that setup to customers.

Signing off on one of these means signing off on four kinds of model working in sequence. A voice-to-text model turns what the questioner says into text. A language model works out a response. A text-to-voice model turns that response into audio, and Synthesia's video model animates the face [13]. The reporter's avatar ran on Synthesia's own video and voice models, but customers can pick alternatives from Cartesia, ElevenLabs, Google or OpenAI [11]. They can also host on a cloud of their choice or pay Synthesia to host [12].

The reporter describes her interactive avatar as "deterministic, meaning it will only say what it was trained to respond to" [14]. In the same piece, the language model in the stack is described as agentic, one that "makes sense of text and can take actions based on it" [13]. To whoever signs off, those are two different promises. What the article does show is scope: both avatars on record answer only from fixed material [1][7].

Capturing a person takes very little. The reporter sat in a small studio inside Synthesia's office for a set of photos and a two-minute voice recording, and she had to consent before anything was made [8]. A couple of days later she had two script-reading versions of herself, with and without glasses, and two interactive ones, four in all [9][10][17].

For a comms or enablement lead, any proposed use can be sorted with two questions. First, who hears it: employees, or people outside the company. Second, what it answers from: a named document, or whatever the model makes of the question. Inside and closed covers the training video Synthesia already sells. Inside and open covers roleplay scoring, where a bad answer lands on an employee practising a pitch [4]. Outside and closed covers Voica's press FAQ and the reporter's single-story avatar. Outside and open is an avatar that takes any question on the company's behalf, and a policy has to deal with that box first.

I'd approve outside-and-closed uses with a named owner. I'd keep the fourth box shut until a vendor shows its avatar stays inside its material when a questioner pushes. The cost of drawing the line there is that a closed avatar only saves time on questions the team saw coming, such as what the company does and how it works [1]. The reporter with the difficult question still ends up with a person.

Any proposal should be able to supply two names. One is the document the avatar answers from. The other is the person who consented to lend it a face and voice [8] and who signs off on what it says. A proposal that cannot name both stays in the training box.

What to watch

  • Whether Synthesia starts selling a press or spokesperson avatar to customers, and on what terms.
  • A published test of what the interactive avatar does when a reporter asks something outside its trained material.
  • Another company's comms team putting an executive's avatar in front of reporters.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories