Skip to content
AI EmployeeSuper-AgentsAgent-to-AgentTutorialsPricingBlogContact

Tavus Griffin and the 48% Video Turing Test, Explained

At a glanceQuick answers
What is Tavus Griffin?
A full-duplex video-to-video model that listens, watches and responds on a live call in one system, generating the whole frame from one reference image. Tavus calls it a Human Interaction Model.
Did it pass a Turing test?
In Tavus’s own one-minute study, 48% of 54 participants said their partner was a real person. That is a company-run test with a small sample, not an independent result.
Can I use it?
Not yet. Griffin-Lite is a research preview for select testers; Tavus says customer release waits on safety and disclosure work.
Editorial data illustration on a near-white ground titled The video Turing test: two video-call windows, Griffin-Lite with 26 of 54 squares filled amber under 48 percent, and Tavus's previous system with 1 of 41 squares filled teal under 2.4 percent
Fig 026 of 54 versus 1 of 41. Made by CellCog's image agent, running GPT Image 2.5.

Tavus says nearly half the people who spoke to its new model on a live video call thought it was a person. On October 1, 2026, the company introduced Griffin, which it calls the first Human Interaction Model, and claimed a milestone in one line: “Griffin is the first model to pass the real-time, video Turing test.” The evidence is a one-minute study Tavus ran itself, plus a top score on NVIDIA’s VideoFDB benchmark.

This page is read from Tavus’s Griffin page, Tavus’s launch post on X and NVIDIA’s VideoFDB leaderboard. The study numbers are the company’s own.

On this page · 8 sectionsOpen
  1. What Griffin is
  2. The study behind the 48%
  3. The benchmark
  4. What is verified and what is not
  5. Why it matters
  6. What we are watching for
  7. Update log
  8. Sources
Key points5 · 6 min full read
  1. On October 1, 2026, Tavus introduced Griffin, a real-time video model it calls the first Human Interaction Model, and claimed it is the first to pass the video Turing test.
  2. In Tavus’s own study, 26 of 54 people who had a one-minute video call with Griffin-Lite said their partner was a real person, against 1 of 41 for Tavus’s previous system.
  3. Participants were told they were talking to another participant, the recruiting platform is unnamed, and no outside evaluator ran the study.
  4. On NVIDIA’s VideoFDB leaderboard, Griffin Lite leads the perception track at 3.73 out of 5, against 4.20 for humans; its median response delay is 2,232 ms, slower than the open-source MiniCPM-o.
  5. Griffin-Lite is a research preview for select testers only. Tavus is withholding customer access until it ships disclosure features, and has published no price.

§ 01What Griffin is

Most video agents chain separate systems: speech recognition, a language model, speech synthesis, then an avatar. Griffin folds them into one full-duplex system that listens and watches while it talks, so it can nod, interrupt, laugh or wait without handing off between models.

Item Detail
Maker Tavus
Announced October 1, 2026, 16:59 UTC on X
What ships Griffin-Lite, a research preview for select testers
Customer access None yet; Tavus says it is waiting on safety work
Video 720p, generated in 320 ms chunks from one reference image
Voice Cloned from about 10 seconds of audio
Audio-to-video delay 0.43 seconds on average on H100s, by Tavus’s measure
Price Not published
Table 1Griffin-Lite at launch, from Tavus’s page

Tavus says the model generates the whole frame, not only the face: the chair, the shadows and the background move too.

§ 02The study behind the 48%

Tavus recruited participants through a research platform it does not name and told them they would talk with another participant for one minute about what they were looking forward to that year. The partner was Griffin-Lite. “Of the 54 participants who spoke with Griffin-Lite, 26 believed their partner was a real person.”

System Participants Said it was a real person Share
Griffin-Lite 54 26 48%
Phoenix-4.5 + Sparrow-2 + Raven-1 (Tavus’s previous stack) 41 1 2.4%
Table 2The face-to-face study, per Tavus
Share who said their one-minute call partner was a real person, per TavusBar chart of Tavus's study: Griffin-Lite 48 percent highlighted against Tavus's previous system at 2.4 percentGriffin-Lite48Previous Tavus system2.4Share who said their one-minute call partner was a real person, per TavusBar chart of Tavus's study: Griffin-Lite 48 percent highlighted against Tavus's previous system at 2.4 percentGriffin-Lite48Previous Tavus system2.4
Fig 1Share who said their one-minute call partner was a real person, per Tavus

Some details matter more than the headline. Those who said the partner was real averaged 79% confidence; those who said AI averaged 81%. People who suspected tended to suspect within the first 20 seconds. On a 7-point scale, participants gave the conversation 4.9 for flowing naturally, its lowest rating.

§ 03The benchmark

VideoFDB is an NVIDIA research benchmark with 237 clips from real video calls. It scores how a model reads a person and how it responds. Tavus says “NVIDIA conducts the evaluation independently using the published metrics and its own judge.” On NVIDIA’s own leaderboard, Griffin-Lite leads the perception track.

System Overall Timing alignment / median latency
Human reference 4.20 90% / 1,400 ms
Tavus Griffin Lite 3.73 73.8% / 2,232 ms
MiniCPM-o 4.5 (open source) 3.40 73% / 720 ms
Gemini 2.5 Flash Native 3.17 72% / 3,160 ms
OpenAI gpt-realtime 2.75 72% / 5,400 ms
Table 3VideoFDB perception track, overall score (0 to 5), NVIDIA leaderboard
VideoFDB perception track, overall score out of 5, from NVIDIA's leaderboardBar chart of VideoFDB perception scores: human reference 4.20, Griffin Lite 3.73 highlighted, MiniCPM-o 4.5 3.40, Gemini 2.5 Flash Native 3.17, OpenAI gpt-realtime 2.75Human reference4.20Griffin Lite3.73MiniCPM-o 4.53.40Gemini 2.5 Flash Native3.17OpenAI gpt-realtime2.75VideoFDB perception track, overall score out of 5, from NVIDIA's leaderboardBar chart of VideoFDB perception scores: human reference 4.20, Griffin Lite 3.73 highlighted, MiniCPM-o 4.5 3.40, Gemini 2.5 Flash Native 3.17, OpenAI gpt-realtime 2.75Human reference4.20Griffin Lite3.73MiniCPM-o 4.53.40Gemini 2.5 Flash Native3.17OpenAI gpt-realtime2.75
Fig 2VideoFDB perception track, overall score out of 5, from NVIDIA's leaderboard

Tavus also reports 3.83 on the generation track, against 3.92 for the human reference and 2.80 for the next system. The timing column is the honest counterweight: Griffin’s median response delay on the perception track is 2,232 ms, behind the open-source MiniCPM-o at 720 ms and the human reference at 1,400 ms.

§ 04What is verified and what is not

  • The benchmark score is on NVIDIA’s page. Griffin Lite sits at the top of the perception leaderboard NVIDIA publishes.
  • The 48% is Tavus’s own study. 54 people, one-minute calls, a recruiting platform it does not name, and no outside evaluator.
  • The setup primed belief. Participants were told their partner was another participant. Tavus says only at the end were they asked “whether it had crossed their mind that their partner might not be a real person.”
  • The comparison number moves. The study section says 2.4% for the previous system; the page’s summary says “a max 2% pass rate” and the X post says under 3%.
  • No paper or price. Tavus cites a Tavus Research write-up; we found no separate paper, and Griffin is not on its platform.

§ 05Why it matters

Tavus names the risk itself: “The same properties that make Human Interaction Models powerful interfaces for natural communications between human and machine allow them to deceive a human into believing it is not AI.” It is holding the model back and says it is building disclosure features first. “Griffin-Lite will not be available for use for customers at this time.”

That is the right order. A model that passes as a person for a minute is a support, tutoring or sales tool only if the person on the other side knows what they are talking to. The benchmark says machines are getting better at reading a face; the study says one minute is now too short to tell. Disclosure has to come from the product, not the viewer.

§ 06What we are watching for

  • A customer release of Griffin, its price, and the disclosure features Tavus promised.
  • An independent replication of the video Turing result with a larger sample and longer calls.
  • A paper with the full method.

§ 07Update log

  • October 1, 2026: page opened, from Tavus’s launch page, its X post and NVIDIA’s leaderboard.

§ 08Sources

Frequently asked5 questions

Q1Who ran the Griffin video Turing test?

Tavus did. It recruited participants through a research platform it does not name, told them they would talk to another participant for one minute, and asked at the end whether they thought the partner was a real person.

Q2Is the NVIDIA benchmark independent?

VideoFDB is NVIDIA’s research benchmark, and NVIDIA’s own leaderboard lists Griffin Lite first on the perception track. Tavus says NVIDIA ran the evaluation with its own judge.

Q3How much does Griffin cost?

Tavus has published no price. Griffin is not on its platform; its existing Phoenix, Raven and Sparrow models remain the ones developers can build with.

Q4Why is Tavus holding it back?

Tavus says the same properties that make the model natural let it deceive people into thinking it is not AI, and that it is building disclosure features and working with safety organizations before a wider release.

Q5Where does CellCog fit?

We build AI employees for any role. They join meetings and calls by voice under their own name and say they are AI. We do not use Tavus, and our employees do not appear on camera today.

Published 02 October 2026 All Trust, permissions & security →