Tavus Griffin passes video Turing test
TECH

Tavus Griffin passes video Turing test

45+
Signals

Strategic Overview

  • 01.
    Tavus introduced Griffin on October 1, 2026, calling it the first 'Human Interaction Model' (HIM) - an AI system built to see, hear, and respond to live video conversation by unifying perception, conversation modeling, and video generation into a single video-to-video pipeline.
  • 02.
    Griffin is full-duplex, meaning it listens, watches, and talks simultaneously rather than waiting for strict turn-taking - letting it backchannel with sounds like 'mm-hm,' tolerate interruptions, and react to visual cues in real time.
  • 03.
    In a Tavus-run study, 26 of 54 participants (48%) believed Griffin-Lite was a real human after a one-minute video call, compared to just 1 of 41 (2.4%) for Tavus's prior Phoenix-4.5 stack.
  • 04.
    Participants were told only that they would be video-chatting with another study participant about what they were looking forward to that year; they weren't asked whether their partner might be AI until after the call ended.
  • 05.
    Griffin-Lite ranked #1 on NVIDIA's independent VideoFDB full-duplex video benchmark, scoring 3.83/5 on generation (versus 2.80 for the next-best system and 3.92 for the human reference) and 3.73/5 on perception.
  • 06.
    The model runs with an average audio-to-video latency of about 0.43 seconds on NVIDIA H100 chips, streaming 720p video in 320-millisecond chunks via an autoregressive diffusion architecture.
  • 07.
    Griffin-Lite remains a research preview limited to select trusted testers; Tavus says it is not yet available to customers while the company builds disclosure and safety mechanisms ahead of a wider release.
  • 08.
    Tavus itself acknowledges the dual-use risk: the same realism that makes Griffin a natural interface for human-machine communication could also let it deceive someone into thinking it isn't AI.
  • 09.
    Tavus frames the same realism as a feature for benign use cases too, citing tutoring, helping people practice difficult conversations, and camera-based tech support as intended applications for Human Interaction Models.

Deep Analysis

From Pipeline to Unified Video-to-Video: How Griffin Achieves Full-Duplex Realism

Tavus built Griffin to replace the old assembly line of separate speech-recognition, language-model, speech-synthesis, and avatar-rendering components with a single unified video-to-video system [1]. That collapse is what makes full-duplex behavior possible - instead of waiting for a speaker to finish before generating a reply, Griffin continuously reassesses the conversation several times a second, letting it backchannel, tolerate interruptions, and track silence duration like a real phone call [2]. Tavus's own launch materials lean into that self-awareness rather than hide it: its official demo video stages the reveal as a video call where one participant turns out to be the AI, with dialogue that jokes about 'passing a Turing test soon' and a 'stop trying to explain it, just show them' pitch for the whole launch, while a more conventional BusinessWire press-release video covered the same announcement in formal terms. The naturalness shows up in smaller, unscripted moments too - one early tester described hopping on a call with Griffin while running a fever, and having the model respond 'oh no kai, a fever is the worst' and ask whether he'd taken anything, the kind of contextually grounded reaction a chained pipeline of separate components struggles to improvise. The payoff also shows up in the harder numbers - Griffin-Lite streams 720p video in 320-millisecond chunks via a streaming autoregressive diffusion model, averaging 0.43 seconds of audio-to-video latency on NVIDIA H100 chips [1][4]. By that measure it looks slower than Tavus's own prior model, Phoenix-4.5, which cited 134ms latency and was marketed as the fastest in market [7]- though the two numbers may not be measuring quite the same thing, since Griffin's figure describes end-to-end audio-to-video response time while Phoenix-4.5's describes per-frame rendering latency. Even allowing for that caveat, Phoenix-4.5 fooled only 2.4% of participants in a comparable study, so raw speed alone clearly isn't what explains believability; conversational timing and multimodal coherence appear to matter more.

What 'Passing a Video Turing Test' Actually Means - and What It Doesn't

The 'video Turing test' framing deserves scrutiny before it gets repeated as fact. Participants in Tavus's study weren't told an AI might be on the call - they were simply informed they'd be matched with another person for a chat about the year ahead, and only asked afterward whether it had crossed their mind their partner wasn't real [2]. That's a meaningfully different setup from a classic Turing test, where evaluators actively try to detect the machine. Participants also rated Griffin 5.4 out of 7 for seeming natural and 5.6 out of 7 for trustworthy - respectable, but well short of a confident pass, and cellcog.ai separately reported a low end of 4.9 out of 7 on naturalness [3][8]. The confidence numbers cut both ways too: participants who believed they'd spoken to a human averaged 79% confidence in that belief, while those who correctly guessed AI averaged 81% confidence, and when suspicion did arise it typically surfaced within the first 20 seconds of the call [1][3]- suggesting the model tends to win over doubters almost immediately or not at all. Cellcog.ai's independent analysis goes further, noting the 48% figure comes from a company-run test with a small, unnamed-platform sample, no external evaluator, and no published paper [3]. X's own community-notes contributors flagged the same gaps, calling the result unverified [2], and YouTube coverage followed a similar pattern - independent channels covering the launch, including one run by a creator known as Prompt Engineer 48, largely repeated Tavus's self-reported 48% figure rather than independently verifying it. None of this means Griffin isn't impressive - but it does mean the headline stat describes a primed, unblinded encounter rather than a rigorously adversarial test, a distinction reflected in a public reaction that split into more than two camps. On X, amplification was mostly enthusiastic, including from AI-news aggregators who pushed the figure to a wider audience. Reddit split further still: r/accelerate read it as impressive but unsettling, with one comment framing the pace of progress as 'one week equals one year' of AI advancement while also noting the demo still feels stiff, with repetitive, PR-style speech patterns; r/nextfuckinglevel was openly hostile, dismissing it as 'AI slop,' alleging astroturfing, and arguing the demo actually fails a Turing test because of visible lag and unnatural responses; and in r/Popular_Science_Ru, viewers catalogued specific rendering tells frame by frame - extra fingers, irregular breathing, lip desync. Older skeptics, meanwhile, pointed out that chatbots like ELIZA already fooled roughly half of users decades before video was even part of the equation.

The Deepfake Question: Why Griffin Is Still Gated

Tavus doesn't hide the trade-off it's created. The company states plainly that the same properties making Human Interaction Models powerful for natural human-machine communication also allow them to deceive someone into believing they aren't talking to AI [1][3]. That's why Griffin-Lite remains a research preview restricted to select trusted testers rather than a shipped product - Tavus says it's building disclosure features and working with safety organizations before any broader release [1][6]. The caution tracks with real precedent: North Korean-linked actors have already used deepfakes on Zoom and Teams calls to target crypto workers, and a model that can generate a convincing live video persona from a single reference image and short voice clip raises the stakes for executive impersonation and interview fraud [2]. Tavus's own pitch for Griffin leans on the benign side of that same realism, too - the company points to use cases like tutoring, helping people rehearse difficult conversations, and camera-based tech support as the kind of natural, trust-building interactions Human Interaction Models are meant to enable, even as it concedes the harder cases need more safeguards before a wider release.

Benchmarked Against NVIDIA and Rivals: Where Griffin Actually Ranks

Stripped of the marketing framing, Griffin's strongest evidence may be its third-party benchmark showing, not its self-run user study. On NVIDIA's independent VideoFDB full-duplex video benchmark, Griffin-Lite scored 3.83 out of 5 on the generation track against 2.80 for the next-best published system and 3.92 for the human reference [1][5]. On the perception track it scored 3.73, ahead of MiniCPM-o 4.5 (3.44), Gemini 2.5 Flash (3.17), and OpenAI's gpt-realtime (2.97), though still short of the 4.20 human baseline [5]. That ranking - against named, comparable rivals on a benchmark Tavus didn't design - is arguably a cleaner signal of where Griffin actually sits in the full-duplex video field than the 48% figure driving the headlines.

Historical Context

2020
Tavus was founded, initially focused on personalized AI videos for sales and marketing before expanding into live conversational video personas.
2025-01
North Korean-linked actors used deepfakes on Zoom and Teams video calls to target crypto workers, a precedent commentators cite when discussing the fraud risk of Griffin's realism.
2026
Tavus's prior real-time rendering model (paired with Sparrow-2 and Raven-1) achieved 134ms latency, which the company called the fastest in market, but only fooled 2.4% of participants in a comparable study.
2026-09
Griffin was scored on NVIDIA's independent VideoFDB full-duplex video benchmark, ahead of its October public unveiling.
2026-10-01
Tavus publicly unveiled Griffin as a research preview (Griffin-Lite), stating it is the first model to pass a video Turing test.

Power Map

Key Players
Subject

Tavus Griffin passes video Turing test

TA

Tavus

San Francisco AI startup (founded 2020, roughly $64 million raised) that built and publicly announced Griffin, controlling access through a trusted-tester research preview and setting the pace of safety disclosure.

HA

Hassaan Raza

Tavus co-founder and CEO who publicly framed Griffin's goal as making conversational computing interfaces 'invisible.'

NV

NVIDIA

Ran the independent VideoFDB full-duplex video benchmark on which Griffin ranked #1, providing third-party benchmark validation for Tavus's performance claims.

ST

Study participants (54 total)

Human subjects in Tavus's own one-minute video-call study whose post-call responses produced the 48% figure central to the Turing-test claim.

Fact Check

8 cited
  1. [1] Tavus Griffin - the first Human Interaction Model
  2. [2] Tavus' Griffin AI Fools People Into Thinking It's Human During Video Calls
  3. [3] Tavus Griffin: A Critical Look at the Video Turing Test Claim
  4. [4] Tavus Griffin AI model
  5. [5] Tavus Unveils Griffin AI Model for Real-Time Human-Like Video Conversations
  6. [6] Tavus Griffin: Specs, Benchmarks, Pricing
  7. [7] Tavus Phoenix-4.5
  8. [8] Tavus Launches Griffin, a Real-Time Video Model That Passed a Video Turing Test

Source Articles

Top 5

THE SIGNAL.

Analysts

“Flagged that Griffin's results come from a company-run test with a small sample size that hasn't been independently verified, and that the study didn't follow standard Turing-test protocol since participants weren't told an AI might be present.”

X community-notes contributors
Public fact-checkers on X

“Cautioned that the 48% figure comes from a company-run test with a small sample, an unnamed recruiting platform, no external evaluator, and no published paper - with participants primed to expect a human on the other end of the call, all of which limits how much weight the claim should carry.”

Cellcog.ai
Independent AI industry blog
The Crowd

“Introducing Griffin, the first model to pass the video Turing test. 48% of people who talked to it live thought it was a real human. Previous systems have had a pass rate <3%. It is #1 on NVIDIA's benchmark for full-duplex AI video. It's the first Human Interaction Model (HIM).”

@@tavus37726

“Face-to-face used to be the one thing AI couldn't fake. Sequoia backed team Tavus launched the first Human Interaction Model (HIM) Griffin that 48% of people who talked to it live thought it was a real human. On their benchmar test, half the people who met this AI on video...”

@@rohanpaul_ai111

“ok so Tavus gave me early access to Griffin and i hopped on a video call with it while running a fever first thing i say is i'm kinda sick. she goes "oh no kai, a fever is the worst" and asks if i took anything then i ask if she watched "the last of us" and the answer "the last...”

@@0xNotFounder95

“Tavus Unveils Griffin AI That Passes Video Turing Test”

@u/EuphoricTomorrow2440129
Broadcast
48% of People Thought This AI Was a Real Human | Introducing Griffin

48% of People Thought This AI Was a Real Human | Introducing Griffin

Griffin by Tavus: The First AI to Pass the Video Turing Test (48% Fooled)

Griffin by Tavus: The First AI to Pass the Video Turing Test (48% Fooled)

Tavus Introduces Griffin, the First Face-to-Face Human Interaction Model, Unlocking the Future...

Tavus Introduces Griffin, the First Face-to-Face Human Interaction Model, Unlocking the Future...

Tavus Griffin passes video Turing test — AI News | Agentic Brew