On October 1, 2026, US-based Tavus announced Griffin, an AI that holds face-to-face-style conversations using voice and video. In the company's experiment, 26 of 54 people who spoke for one minute with Griffin-Lite, a research-preview version, judged that they were talking to a real person. Griffin is designed to keep seeing and hearing the other person while it speaks, and to return backchannel responses and facial expressions in the middle of a conversation. The aim is to handle not only a human-looking face and voice but also the "timing" of a conversation. Still, the share of people who thought it was human, the ability to respond naturally, and the speed of reaction are separate metrics, and each needs to be checked on its own.
What happened in the one-minute in-house experiment
The 26-of-54 result depends on what participants were told before the call. Tavus recruited participants through an external research platform and told them they would talk for one minute with "another participant" about what they were looking forward to this year. The actual counterpart was Griffin-Lite, which generated its face and voice on the spot. This was not a test in which participants were assigned the task of telling AI from human in advance.
After the call, participants wrote down what the other party had said and rated naturalness and conversational flow. At the end of the survey, they were asked whether they thought the other party might not be human, and finally everyone was told it had been an AI. With the earlier system, which combined Phoenix-4.5, Sparrow-2, and Raven-1, the same procedure produced 1 of 41 people, about 2.4%, who judged the counterpart to be human. This time the figure was about 48%. Compared with the company's previous system, at least, the cases in which the AI could hold a conversation while keeping the impression of being human increased substantially.
Tavus describes this as "the first model to pass the video Turing test." However, the published experimental method and results do not show the performance of a control group that spoke with real humans under the same conditions, nor whether participants were randomly assigned. Recruiting through an external platform also does not mean a third party ran the experiment independently. The 48% figure cannot simply be converted into the conclusion that telling humans and AI apart has become as hard as a coin toss.
The test lasted one minute and the topic was limited. This experiment alone cannot tell us whether the same result would hold in long conversations or in situations where people are told from the start that the other party is an AI. According to Tavus, people who suspected an AI tended to feel so within the first 20 seconds. Even in a short call, some people sensed something was off; Griffin has not reached a stage where it can make anyone believe it is human.
Returning backchannels and expressions while still talking
When a speaker stops talking and looks away, that silence can signal either "I'm finished" or "I'm still thinking." VideoFDB, an evaluation method developed by NVIDIA and David AI, raises this distinction as a challenge for conversational AI. Even when a break in the audio suggests it is fine to begin replying, taking gaze into account may show that waiting is the right call. Being able to describe what is on screen is a different ability from using that footage as a cue to judge conversational timing.
The conventional approach, as Tavus describes it, chains processes in sequence: speech is converted to text, a language model writes a reply, and that content is converted into voice and video. In this configuration, the later video-generation stage moves in response to the spoken audio, which makes it hard to build in nods while the other person is speaking or expression changes made silently. VideoFDB research likewise reports that the audio-driven avatars it evaluated lacked these nonverbal responses.
Griffin runs conversational decision-making and audio/video generation in parallel. The part responsible for judgment receives audio and video and repeatedly decides whether to speak or wait, at intervals of under one second. The control signals it sends include not only what to say but also instructions for tone of voice and gestures. The generation side follows those instructions, producing voice and video bit by bit, and keeps seeing and hearing the other person even during playback. It is a "full-duplex" dialogue, listening while speaking.
However, the word "integrated" does not mean everything is processed by a single neural network. Tavus's own technical explanation shows a structure split into a part that judges the conversation and a part that generates audio and video. The major design change is that each part does not wait for the other person to finish speaking and then act in sequence, but keeps operating continuously, including reactions made mid-utterance.
On the video side, the company says it generates not just the face but the body and background from a single reference image. It outputs 720p video at 25 frames per second, generating 320 milliseconds of video, or 8 frames, at a time. The audio side also outputs without waiting for a whole sentence to be completed, and can play back in units as small as 10 milliseconds. Generating in such fine units makes it possible to reflect tone and motion instructions that arrive midway into subsequent output.
The 320 milliseconds here is the length of video generated at one time, not a response time to a person's speech. The 0.43-second figure Tavus published separately is the average delay on an H100 between feeding audio into the video generator and the effect of that audio appearing in the video. The response speed including the processing that decides conversational content needs to be looked at in a separate evaluation.
Scores close to human, and a remaining timing gap
In NVIDIA's published table, Griffin-Lite scored 3.83 in the VideoFDB generation category. The human reference value was 3.92, and the runner-up, a configuration combining Gemini 2.5 and Anam, scored 2.80. This is not the share of people who thought the counterpart was human, but a value from scoring the audio and video the AI returned against evaluation criteria. VideoFDB uses 237 clips collected from real video calls and is scored from 0 to 5 by a language model.
In the perception category, which measures reading the other party's behavior and connecting it to an appropriate reply, Griffin-Lite also leads the published table with 3.73. The human reference value was 4.20. The highest score among the comparison targets, 3.44, was for MiniCPM-o 4.5 configured without video input. Because the generation and perception categories evaluate different abilities and different comparison targets, they cannot simply be combined into one overall ranking.
Looking at the 3.83 generation score item by item shows where Griffin-Lite has approached humans and where a gap remains.
| VideoFDB generation metric | Griffin-Lite | Human reference |
|---|---|---|
| Fluency | 4.25 | 4.42 |
| Emotional alignment with the other party | 4.40 | 4.14 |
| Appropriateness of nonverbal behavior | 2.83 | 3.18 |
| Overall score | 3.83 | 3.92 |
| Speaking-timing match rate | 62.8% | 78% |
| Median latency | 1,892 ms | 900 ms |
The source is NVIDIA's generation-category table, checked on October 2, 2026. Scores are on a 0–5 scale; timing match rate and latency are separate metrics.
Griffin-Lite's emotional expression exceeds the human reference value. On the other hand, the nonverbal-behavior rating, which measures whether expressions and gestures suit the situation, is lower than the human value. Being good at conveying emotion through voice and expression does not guarantee delivering it at the right moment. Taking an overall score close to human as meaning every kind of reaction is now human-level would overlook this difference.
In NVIDIA's VideoFDB generation category, Griffin-Lite's median latency was 1,892 milliseconds, about 2.10 times the human reference of 900 milliseconds. This figure is calculated as 1,892 ÷ 900, comparing the same metric in the same category. However, the latency defined in Appendix C.2 of the paper is based on events annotated during the conversation. It measures situations such as starting to speak or yielding the floor after being interrupted, and should be distinguished from the time between receiving a question and replying, and from the video generator's standalone average of 0.43 seconds.
The speaking-timing match rate was also 62.8% for Griffin-Lite against 78% for humans. There are moments when starting to speak immediately is appropriate and moments when waiting silently is more natural. Griffin's progress shows in the gap from avatars built on conventional sequential processing, but the challenge of choosing the right "beat" remains.
The paper by Amrita Mazumdar and colleagues, who proposed VideoFDB, is a preprint published in May 2026. Appendix B states that the scope is English-language two-person video calls in the US and Canada, and that evaluation targets a single turn within a conversation. Although the latest rankings can be checked in NVIDIA's public table, they are not results demonstrating Japanese-language conversation or the building of relationships over a long period.
Seeming human versus being easy to talk to
In Tavus's call experiment, even participants who guessed the counterpart was an AI gave the "want to talk again" item a 5.4 on a 7-point scale. The overall average was 5.8. This shows that even participants who suspected an AI were willing to converse again. However, because there was no comparison experiment in which people were told from the start that it was an AI, the figure does not indicate how readily people would accept it when its AI nature is disclosed.
Uses the company envisions include learning support that reads signs of confusion from facial expressions and changes how it explains, and practice for difficult conversations. It also envisions support scenarios in which a user shows a broken part to the camera while asking for help. In each case, what helps is less that the face looks real than that the AI can read how much the other person understands and when they need a reply. At present, though, the effectiveness of such uses has not been demonstrated; they remain examples of future use that Tavus has presented.
Rollout is also being handled cautiously. Griffin-Lite is limited to a research preview for some testers and is not yet available to general customers. Tavus says it is developing safeguards, including features that tell the other party they are talking to an AI, and plans to release a higher-performance version after addressing such issues. It has not disclosed a specific general availability date or pricing.
The existing Tavus acceptable use policy requires explicit consent from the person whose face or voice is replicated, and obliges users to make clear to conversation partners that they are speaking with an AI. A person permitting their face or voice to be replicated and a conversation partner knowing who they are talking to are separate conditions, each of which must be met. However, the policy requiring disclosure is not the same as a mechanism being implemented that automatically tells people it is an AI on every call using Griffin.
What Griffin shows is a change in evaluating conversational AI: it is no longer enough to attach a face to the text of a reply; motion and the "beats" of conversation also need to be examined. Even when users understand they are dealing with an AI, if it can respect their time to think and adjust its explanations to their confusion, it can reduce the burden on humans of changing how they speak to suit a machine. Whether it can truly support learning and conversation practice under those conditions will be the basis for judging whether human-like video can be turned into practical dialogue.
