Build4 publishers3 min readPublished Updated
Tavus's Griffin video model passed for human with 26 of 54 people in a company-run test
Tavus says 26 of 54 people took its Griffin video model for a real person in one-minute live calls, a 48% rate from its own study. The company is keeping Griffin from customers until its disclosure and safety features are ready.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- Participants in the US and Europe were told they would spend a minute talking with another participant about the year ahead, and were asked only at the end whether their partner was real.
- Under the same protocol, one of 41 participants, or 2.4%, mistook Tavus's previous Phoenix-4.5 system for a person.
- Tavus says NVIDIA scored Griffin-Lite on its Video Full-Duplex Benchmark in September and ranked it first on both the generation and perception tracks.
- Tavus announced Griffin-Lite on X on October 1 as a research preview open only to selected testers.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint On Tavus's own benchmark scores, Griffin is furthest from human at reading a conversation, and reading what the camera shows is the skill a camera-based support agent needs most.
- decision Teams building on Tavus today are building on its existing conversational video tier, because Griffin sits outside the paid plans.
- precedent With Tavus setting working disclosure features as its own release bar, buyers of rival real-time video agents have a concrete question for each vendor: how does the persona tell callers it is software?
Because the partner was billed as a fellow participant, the setup leaned answers toward "human" before anyone spoke [3]. Even so, 28 of the 54 did not take Griffin for a person [2]. The Phoenix-4.5 arm controls for that lean, since the same cover story produced almost no human verdicts for the older system [4]. The published account does not say whether participants were assigned to the two arms at random or tested in the same period.
Fifty-four people is a small sample. A normal-approximation 95% interval on 26 of 54 runs from about 35% to 62% [3]. RuntimeWire, reporting the announcement, describes it as a small company-run result and not a broad measure of how often people would mistake Griffin for a human in ordinary use [13]. For the number to transfer, a deployment would have to look like the test: one minute, small talk about the year ahead, and a caller with no reason to suspect software. Tavus lists tutoring, practice for difficult conversations and camera-based tech support as uses [16]. Each of those runs past a minute and keeps the model on a task.
The design target is the right one. Griffin takes in audio and video continuously and decides whether to speak, wait or yield, producing voice and video together [9]. Building the model around turn-taking is sound engineering, because turn-taking is what a one-minute call exposes first. Tavus says it renders full scenes and responds to objects held up to the camera, and one demo has it playing Simon Says [10]. That is a fair turn-taking test, if not one a buyer will write into a contract.
The latency figure covers one component. Tavus says the video-generation stage answers incoming audio with a frame in an average of 0.43 seconds on NVIDIA H100s, about half the latency of the next-fastest published system it compared against [11]. A caller notices the slowest replies, and a mean on top-end GPUs describes the typical one.
The NVIDIA scores come from a language-model judge, and RuntimeWire notes they measure something different from the live study [8]. Tavus describes the benchmark run as independent, according to The Decoder [14]. Griffin-Lite scored 3.83 of 5 on generation against a human reference of 3.92, and 3.73 on perception against 4.20 [6][7]. That leaves it 0.09 points short of human on generation and 0.47 short on perception [4][5]. Its lead over the next system is 1.03 points on generation and 0.29 on perception [6][7].
Tavus is withholding Griffin because the realism that makes it useful can also make it deceptive [12]. A more capable version will follow once "safety concerns are addressed," the company says [15]. The disclosure gate is Tavus's own decision. I think it is the right one for a company whose own test had about half its participants fooled inside a minute [1]. Tavus raised a $40 million Series B in November 2025, led by CRV [17], and about $64 million in total, according to The Decoder [18].
What to watch
- Tavus's full research report: whether participants were randomly assigned between the Griffin and Phoenix-4.5 arms, and whether longer, task-bound calls keep the rate near half.
- The disclosure features Tavus ships with Griffin, and whether the model then appears on its paid conversational video plans.
- Tail or end-to-end latency figures for Griffin beyond the 0.43-second average for the video-generation component.