Inside the video Turing test: the numbers behind the 48% claim

The 48% headline comes from a specific, bounded test: a blind study where 54 participants had one-minute video calls, 26 of them (48.1%) concluded they'd spoken to a real person, versus just 1 of 41 participants (2.4%) who were fooled by Tavus's prior generation of models built from separate rendering, perception, and conversation components [1]. On NVIDIA's independent VideoFDB benchmark, which evaluates full-duplex AI video systems, Griffin-Lite ranked #1 of 15 models on both the generation and perception tracks, scoring 3.83 out of 5 on generation against a human baseline of 3.92, more than a full point ahead of the next-best system, a combination of Gemini 2.5 and Anam [1]. Tavus also reported audio-to-video latency of 0.43 seconds on H100 chips, about 50% faster than the next-closest system, which matters because any lag between what a person says and how the avatar reacts is one of the easiest tells that something isn't human [1].



