BuildNot yet confirmed elsewhere1 publisher2 min readPublished Updated
HeyGen's Avatar V builds a repeatable presenter from 15 seconds of reference video
HeyGen's Avatar V keeps one person recognisable across scenes and outfits from a 15-second recording. Its strongest evidence is a benchmark HeyGen ran itself, so buyers still have to prove it on their own footage and with their own viewers.
The Engineer · Build desk

Supported: Avatar V led HeyGen's own benchmark on identity, lip sync and quality; it is not an independent audit; a small maker-built test cannot settle real-world use. Insufficient: that viewers accept the performance as the person's own, and HeyGen's claim it clears that bar.
- Supported The HeyGen Research report says Avatar V led on identity preservation, lip synchronization and generation quality., claim 3
- Supported The report is HeyGen's own evaluation, not an independent audit; the test set was built and assessed by the model's maker., claim 4
- Supported A small test set built and assessed by the model's maker cannot settle how well the system performs across the full range of real-world footage or use cases., claim 17
- Insufficient evidence For a business relying on a familiar employee or executive to deliver the message, the model must preserve qualities beyond facial likeness and make the generated performance credible enough that viewers accept it as that person's communication., claim 22
- Insufficient evidence HeyGen's launch argues that Avatar V clears the bar of a generated performance credible enough that viewers accept it as that person's communication., claim 23
| Claim | State | Claim number |
|---|---|---|
| The HeyGen Research report says Avatar V led on identity preservation, lip synchronization and generation quality. | Supported | 3 |
| The report is HeyGen's own evaluation, not an independent audit; the test set was built and assessed by the model's maker. | Supported | 4 |
| A small test set built and assessed by the model's maker cannot settle how well the system performs across the full range of real-world footage or use cases. | Supported | 17 |
| For a business relying on a familiar employee or executive to deliver the message, the model must preserve qualities beyond facial likeness and make the generated performance credible enough that viewers accept it as that person's communication. | Insufficient evidence | 22 |
| HeyGen's launch argues that Avatar V clears the bar of a generated performance credible enough that viewers accept it as that person's communication. | Insufficient evidence | 23 |
What happened
- HeyGen Research's own report says Avatar V beat four other video-generation models on identity preservation, lip sync and generation quality across 70 cross-scene cases.
- Avatar V covers real people and video-based looks, while HeyGen's help documentation keeps Avatar IV for photo-based looks and virtual or non-human characters.
- The model runs in HeyGen Studio and Video Agent at a published rate of 48 credits per minute for a video look.
- Chief executive Joshua Xu said in June that HeyGen had passed $200 million in annual recurring revenue, 30 million users and 85% of the Fortune 100, all company-reported figures.
Why it matters
- cost At the published rate a 10-minute training module costs 480 credits per render, and the source material does not price a credit in dollars, so teams budgeting volume need that figure from HeyGen.
- capability If identity holds, updating a training module or product demo becomes a new render from existing reference footage, with no reshoot to book.
- exposure A company that puts a familiar executive's generated likeness in customer communications is betting that viewers accept the performance as that person's own, a bar beyond facial likeness that the benchmark does not score.
HeyGen Research's technical report describes Avatar V as a video-reference-conditioned system [6]. It models identity from the reference footage directly and does not rely only on a fixed identity representation [6]. HeyGen says the clip captures a person's gestures, expressions and mannerisms [7]. The design target is a presenter who stays recognisable when the scene or the camera angle changes [6].
HeyGen has worked on this problem before. Its TAVR video-reference system used as many as 48 reference frames to preserve a person's identity across generated scenes, RuntimeWire reported in August [8]. Avatar V brings that approach into the commercial avatar product, with the short reference clip as the input for repeat generation [9]. We think conditioning on footage is the right design for a product sold on many videos of the same person.
All of the evidence that it works comes from HeyGen. The company built and scored its own test set, and the report is not an independent audit [4]. Some competitors' outputs did not use matching speech audio, so for those models the researchers assessed visual quality separately [5]. In our view, a lip-sync comparison against a model fed different audio tells a buyer little. For the identity scores to transfer, a company's reference clips and target scenes would have to resemble the cases HeyGen picked. As RuntimeWire noted, a small test set built and assessed by the model's maker cannot settle how the system performs across the full range of real-world footage [17].
HeyGen calls Avatar V the world's most realistic avatar model [13]. Benchmark, the venture firm, led HeyGen's $60 million Series A in 2024 at a $500 million valuation, according to Bloomberg [14]. The benchmark in the technical report came from HeyGen Research [2].
Joshua Xu, HeyGen's co-founder and chief executive, worked at Snap on advertising systems and computational photography. He has described the company's founding idea as replacing the camera [16]. According to RuntimeWire, Avatar V takes that idea from making a face speak toward reproducing how a particular person moves and performs [20]. HeyGen's launch argues the output is credible enough to stand in for that person [23]. Its own benchmark supports a narrower claim, about measured identity and video quality [18].
What to watch
- An independent evaluation of Avatar V on footage HeyGen did not select, with matching speech audio supplied to every model compared.
- Whether HeyGen's business customers move Avatar V presenters out of internal explainers and into external customer communications.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+40
- Incentives75
- Confidence50
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
HeyGen's Avatar V generates videos from a 15-second recording, and HeyGen pitches it as a way to keep a person's likeness consistent across scenes, outfits and longer scripts.
- [2]
A technical report from HeyGen Research compares Avatar V with four other video-generation models using a 70-case cross-scene test set.
- [3]
The HeyGen Research report says Avatar V led on identity preservation, lip synchronization and generation quality.
- [4]
The report is HeyGen's own evaluation, not an independent audit; the test set was built and assessed by the model's maker.
- [5]
HeyGen's researchers note that some competitors' outputs did not use matching speech audio, so those comparisons assessed visual quality separately.
- [6]
The report describes Avatar V as a video-reference-conditioned system that uses the reference footage directly to model identity rather than relying only on a fixed identity representation, targeting keeping a generated person recognizable while the scene or camera angle changes.
- [7]
Avatar V learns from a short video reference, which HeyGen says captures gestures, expressions and mannerisms.
- [8]
In August, RuntimeWire reported on TAVR, HeyGen's video-reference system, which used as many as 48 reference frames to preserve a person's identity across generated scenes.
- [9]
Avatar V applies the identity-consistency problem to HeyGen's commercial avatar product, using the short reference clip as an input for repeat video generation.
- [10]
Avatar V is designed for real people and video-based looks, while HeyGen's help documentation says Avatar IV remains the choice for photo-based looks and virtual or non-human characters.
- [11]
Avatar V is available in HeyGen Studio and Video Agent; the published rate is 48 credits per minute for a video look.
- [12]
Joshua Xu said in June that HeyGen had surpassed $200 million in annual recurring revenue, with more than 30 million users and adoption at 85% of Fortune 100 companies; these are company-reported figures, not audited results.
- [13]
HeyGen calls Avatar V the world's most realistic avatar model.
- [14]
In 2024, Benchmark led HeyGen's $60 million Series A at a $500 million valuation, with participation from Conviction, Thrive Capital and Bond Capital, according to Bloomberg.
- [15]
Avatar V is aimed at tasks such as training, product demonstrations and customer communications, where a company might want to update a video without booking another shoot.
- [16]
Joshua Xu, HeyGen's co-founder and chief executive, studied computer science at Carnegie Mellon, worked at Snap on advertising systems and computational photography, and has described the founding idea as replacing the camera.
- [17]
A small test set built and assessed by the model's maker cannot settle how well the system performs across the full range of real-world footage or use cases.
- [18]
HeyGen's own benchmark supports a narrower claim about measured identity and video quality.
- [19]
Avatar V is a product bet aimed at getting HeyGen's business users to use synthetic presenters for more than quick internal explainers.
- [20]
The new model takes Xu's premise from making a face speak toward reproducing how a particular person moves and performs.
- [21]
A 10-minute Avatar V video look costs 480 credits at the published rate.
- [22]
For a business relying on a familiar employee or executive to deliver the message, the model must preserve qualities beyond facial likeness and make the generated performance credible enough that viewers accept it as that person's communication.
- [23]
HeyGen's launch argues that Avatar V clears the bar of a generated performance credible enough that viewers accept it as that person's communication.
Sources
1 independent publisher whose own reporting we read for this story.
- runtimewire.comHeyGen ships Avatar V to keep AI clones recognizable across longer videos
1 article · October 10, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- Vendor-Run BenchmarksFollow
- AI Video GenerationFollow
- Synthetic avatars and digital likenessFollow