Tavus says nearly half the people who spoke to its new model on a live video call thought it was a person. On October 1, 2026, the company introduced Griffin, which it calls the first Human Interaction Model, and claimed a milestone in one line: “Griffin is the first model to pass the real-time, video Turing test.” The evidence is a one-minute study Tavus ran itself, plus a top score on NVIDIA’s VideoFDB benchmark.
This page is read from Tavus’s Griffin page, Tavus’s launch post on X and NVIDIA’s VideoFDB leaderboard. The study numbers are the company’s own.
On this page · 8 sectionsOpen
- On October 1, 2026, Tavus introduced Griffin, a real-time video model it calls the first Human Interaction Model, and claimed it is the first to pass the video Turing test.
- In Tavus’s own study, 26 of 54 people who had a one-minute video call with Griffin-Lite said their partner was a real person, against 1 of 41 for Tavus’s previous system.
- Participants were told they were talking to another participant, the recruiting platform is unnamed, and no outside evaluator ran the study.
- On NVIDIA’s VideoFDB leaderboard, Griffin Lite leads the perception track at 3.73 out of 5, against 4.20 for humans; its median response delay is 2,232 ms, slower than the open-source MiniCPM-o.
- Griffin-Lite is a research preview for select testers only. Tavus is withholding customer access until it ships disclosure features, and has published no price.
§ 01What Griffin is
Most video agents chain separate systems: speech recognition, a language model, speech synthesis, then an avatar. Griffin folds them into one full-duplex system that listens and watches while it talks, so it can nod, interrupt, laugh or wait without handing off between models.
| Item | Detail |
|---|---|
| Maker | Tavus |
| Announced | October 1, 2026, 16:59 UTC on X |
| What ships | Griffin-Lite, a research preview for select testers |
| Customer access | None yet; Tavus says it is waiting on safety work |
| Video | 720p, generated in 320 ms chunks from one reference image |
| Voice | Cloned from about 10 seconds of audio |
| Audio-to-video delay | 0.43 seconds on average on H100s, by Tavus’s measure |
| Price | Not published |
Tavus says the model generates the whole frame, not only the face: the chair, the shadows and the background move too.
§ 02The study behind the 48%
Tavus recruited participants through a research platform it does not name and told them they would talk with another participant for one minute about what they were looking forward to that year. The partner was Griffin-Lite. “Of the 54 participants who spoke with Griffin-Lite, 26 believed their partner was a real person.”
| System | Participants | Said it was a real person | Share |
|---|---|---|---|
| Griffin-Lite | 54 | 26 | 48% |
| Phoenix-4.5 + Sparrow-2 + Raven-1 (Tavus’s previous stack) | 41 | 1 | 2.4% |
Some details matter more than the headline. Those who said the partner was real averaged 79% confidence; those who said AI averaged 81%. People who suspected tended to suspect within the first 20 seconds. On a 7-point scale, participants gave the conversation 4.9 for flowing naturally, its lowest rating.
§ 03The benchmark
VideoFDB is an NVIDIA research benchmark with 237 clips from real video calls. It scores how a model reads a person and how it responds. Tavus says “NVIDIA conducts the evaluation independently using the published metrics and its own judge.” On NVIDIA’s own leaderboard, Griffin-Lite leads the perception track.
| System | Overall | Timing alignment / median latency |
|---|---|---|
| Human reference | 4.20 | 90% / 1,400 ms |
| Tavus Griffin Lite | 3.73 | 73.8% / 2,232 ms |
| MiniCPM-o 4.5 (open source) | 3.40 | 73% / 720 ms |
| Gemini 2.5 Flash Native | 3.17 | 72% / 3,160 ms |
| OpenAI gpt-realtime | 2.75 | 72% / 5,400 ms |
Tavus also reports 3.83 on the generation track, against 3.92 for the human reference and 2.80 for the next system. The timing column is the honest counterweight: Griffin’s median response delay on the perception track is 2,232 ms, behind the open-source MiniCPM-o at 720 ms and the human reference at 1,400 ms.
§ 04What is verified and what is not
- The benchmark score is on NVIDIA’s page. Griffin Lite sits at the top of the perception leaderboard NVIDIA publishes.
- The 48% is Tavus’s own study. 54 people, one-minute calls, a recruiting platform it does not name, and no outside evaluator.
- The setup primed belief. Participants were told their partner was another participant. Tavus says only at the end were they asked “whether it had crossed their mind that their partner might not be a real person.”
- The comparison number moves. The study section says 2.4% for the previous system; the page’s summary says “a max 2% pass rate” and the X post says under 3%.
- No paper or price. Tavus cites a Tavus Research write-up; we found no separate paper, and Griffin is not on its platform.
§ 05Why it matters
Tavus names the risk itself: “The same properties that make Human Interaction Models powerful interfaces for natural communications between human and machine allow them to deceive a human into believing it is not AI.” It is holding the model back and says it is building disclosure features first. “Griffin-Lite will not be available for use for customers at this time.”
That is the right order. A model that passes as a person for a minute is a support, tutoring or sales tool only if the person on the other side knows what they are talking to. The benchmark says machines are getting better at reading a face; the study says one minute is now too short to tell. Disclosure has to come from the product, not the viewer.
§ 06What we are watching for
- A customer release of Griffin, its price, and the disclosure features Tavus promised.
- An independent replication of the video Turing result with a larger sample and longer calls.
- A paper with the full method.
§ 07Update log
- October 1, 2026: page opened, from Tavus’s launch page, its X post and NVIDIA’s leaderboard.
§ 08Sources
- Tavus, Griffin: The First Human Interaction Model, October 1, 2026.
- Tavus on X, October 1, 2026.
- NVIDIA, VideoFDB benchmark and leaderboard.
Q1Who ran the Griffin video Turing test?
Tavus did. It recruited participants through a research platform it does not name, told them they would talk to another participant for one minute, and asked at the end whether they thought the partner was a real person.
Q2Is the NVIDIA benchmark independent?
VideoFDB is NVIDIA’s research benchmark, and NVIDIA’s own leaderboard lists Griffin Lite first on the perception track. Tavus says NVIDIA ran the evaluation with its own judge.
Q3How much does Griffin cost?
Tavus has published no price. Griffin is not on its platform; its existing Phoenix, Raven and Sparrow models remain the ones developers can build with.
Q4Why is Tavus holding it back?
Tavus says the same properties that make the model natural let it deceive people into thinking it is not AI, and that it is building disclosure features and working with safety organizations before a wider release.
Q5Where does CellCog fit?
We build AI employees for any role. They join meetings and calls by voice under their own name and say they are AI. We do not use Tavus, and our employees do not appear on camera today.
