On October 1, 2026, Tavus published Griffin and called it the first Human Interaction Model (HIM): one video-to-video, full-duplex system that listens while it talks and watches expressions and pauses, not just transcripts. The research post is signed from San Francisco by co-founder and CEO Hassaan Raza, head of research Ioannis Patras, and the Tavus research team.
The number that traveled is 48%. The number you should keep is 26 of 54 after a one-minute call. The previous in-house stack — Phoenix 4.5 + Sparrow-2 + Raven-1 — was 1 of 41 (2.4%). Marketing said under 3%. The launch thread's lead post passed 7.2 million views. None of that makes Griffin a product you can put in a customer PAL this week.
This is a perception-and-timing story, not a "new avatar skin" story. It sits next to the HeyGen LiveAvatar + GPT-Live-1 demos (cascade you can fork) and the GPT-Live-1 API (audio full-duplex without a generated face). It also sits next to the $25.6 million deepfake video-call fraud pattern, because Tavus delayed customer GA for the same reason the study looks impressive: the model can deceive.
TL;DR
| Question | Direct answer |
|---|---|
| What is HIM? | Human Interaction Model — Tavus's name for a unified video-to-video, full-duplex interaction model. |
| What shipped Oct 1? | A research post plus Griffin-Lite for trusted testers only. Not customer GA. |
| Video Turing headline? | 26/54 (48%) said "real person" after 1 minute. Prior stack: 1/41 (2.4%). |
| VideoFDB generation? | 3.83 vs human 3.92 vs Gemini 2.5 + Anam 2.80 (NVIDIA, Sept 2026). |
| VideoFDB perception? | 3.73 vs MiniCPM-o 4.5 3.44 vs Gemini 2.5 Flash Native 3.17 vs gpt-realtime 2.97. |
| TOR alignment? | 62.8% generation / 73.8% perception. |
| "37% ahead"? | Marketing gloss on a gap. Use the NVIDIA scores, not the percentage. |
| What do existing PALs run? | Still Phoenix / Raven / Sparrow. Griffin is not on the platform yet. |
| Why the delay? | Tavus: further alignment is required because HIMs can deceive. |
What people are asking after the 7.2 million-view thread
The official X thread is seven posts (2105704169009246248, then 2105704243890164166, 2105704311330390313, 2105704437478281436, 2105704541190856704, 2105704783185375402, 2105704914156724693). The questions that follow a thread like that are not "what is a griffin." They are product questions.
Did it really pass a Turing test?
Tavus writes that Griffin is "the first model to pass the real-time, video Turing test." The protocol is specific:
- Participants came through an independent research platform.
- They were told they would be matched with another participant for a one-minute call about what they were looking forward to this year.
- The partner was a PAL on Griffin-Lite, generating face, voice, and replies live. A matched run used Phoenix 4.5 + Sparrow-2 + Raven-1.
- After the call they wrote down the partner's answer and rated naturalness, trust, flow, and whether they felt listened to.
- Only at the end were they asked whether it had crossed their mind that the partner might not be a person. Then everyone was told it was AI.
Results Tavus published:
| Measure | Griffin-Lite (n=54) | Phoenix 4.5 + Sparrow-2 + Raven-1 (n=41) |
|---|---|---|
| Said "real person" | 26 / 54 = 48% | 1 / 41 = 2.4% |
| Confidence if "real" | 79% | — |
| Confidence if "AI" | 81% | — |
| When doubters first suspected | Usually under 20 seconds | — |
| Natural (7-pt) | 5.4 | — |
| Trustworthy (7-pt) | 5.6 | — |
| Would talk again (7-pt) | 5.8 overall, 5.4 among those who said AI | — |
| Felt listened to (7-pt) | 5.5 | — |
| Conversation flowed (7-pt) | 4.9 (lowest of the five) | — |
That is a short, prompted, in-house study, not an hour-long interview with an adversarial interviewer. Over half said the AI possibility never crossed their mind during the call, and nearly all of that group then said the partner was real. People who spotted it usually spotted it fast. The "flow" score is the weak axis. Treat "passed the video Turing test" as Tavus's label for 26/54 on this protocol, not as a standardized exam.
Is Griffin generally available?
No. The same post is explicit: Griffin-Lite "will not be available for use for customers at this time." It is a research preview for select trusted testers. A form exists to request access. Tavus says it expects to release Griffin after safety work, including "safe disclosure features" and work with AI-safety organizations.
The footer is the product map you actually ship against: Griffin is not on the Tavus platform yet. Until it is, you build PALs on the models the company says 150,000 developers and businesses already use: Phoenix generates the face, Raven sees, Sparrow knows when to talk.
Is this just another talking-head model?
No. The research claim is class, not skin. A HIM is designed to understand and generate face-to-face interaction in real time. Perception, turn decisions, and expressive speech/video run together. That is a different object from:
- a cascade (ASR → LLM → TTS → avatar),
- a voice-only full-duplex model such as GPT-Live-1,
- a face layer you bolt onto someone else's brain, which is how LiveAvatar LITE is wired.
Raza's framing in the post: machines have required people to meet them at the prompt. Tavus wants computing that "become[s] invisible." You can take that as a product thesis without taking it as a ship date.
What NVIDIA VideoFDB actually measured
Tavus's page leads with "#1 on NVIDIA's independent test" and "37% ahead of the next best AI at reacting in the moment." Do not treat 37% as the score. NVIDIA's VideoFDB (Mazumdar, Park, Roy, Srihari, Wang, Zhou, Wang, Nagano, De Mello; NVIDIA and David AI; arXiv 2605.30256) is a 0–5 LM-judge benchmark on 237 dyadic clips from real video calls, split into perception and generation, with takeover-rate (TOR) alignment for timing. NVIDIA scored Griffin-Lite in September 2026.
Tavus's published NVIDIA numbers — the ones to quote:
| Track | Griffin-Lite | Human reference | Next systems Tavus cites |
|---|---|---|---|
| Generation (voice + face of the agent) | 3.83 | 3.92 | Gemini 2.5 + Anam 2.80 |
| Perception (does the spoken reply use what it sees?) | 3.73 | 4.20 (human; Tavus) | MiniCPM-o 4.5 3.44, Gemini 2.5 Flash Native 3.17, gpt-realtime 2.97 |
| TOR alignment | 62.8% gen / 73.8% perc | Higher on NVIDIA's human row | Best-in-class in Tavus's write-up |
The generation gap versus Gemini 2.5 + Anam is +1.03 on a 5-point scale (3.83 − 2.80). 1.03 / 2.80 ≈ 37%. That is where the marketing line comes from. It is a relative gap against one cascade, not a universal "37% more human" meter. Griffin is 0.09 below the human generation reference. Tavus also writes it is "more than 12x closer to human performance than any other system evaluated" on that track — again a derived closeness (0.09 vs 1.12), not a second official NVIDIA column.
On perception, Tavus's own caption says Griffin sees audio and video, while MiniCPM-o 4.5 and gpt-realtime are cited in audio-only configurations for those 3.44 and 2.97 figures. NVIDIA's public tables also list Griffin Lite perception 3.73 with TOR 73.8% / 2232 ms, and generation 3.83 with TOR 62.8% / 1892 ms. Use those. Do not flatten them into "37%."
VideoFDB's paper-level finding still matters for builders: cascaded speech-to-avatar pipelines keep turn-yielding discipline but cannot insert nonverbal cues during the user's turn, and they sit 2.8–3.5 seconds behind human ground truth. That is the hole Griffin claims to fill. It is also why a HeyGen-style stack (GPT-Live brain/voice + LiveAvatar face + HyperFrames overlays) is a different product — useful, forkable, and structurally not a HIM.
NVIDIA's benchmark also names failure modes Griffin is trying to dodge: captioning collapse (the model describes your face instead of talking with you) and visual-stream ignorance (audio-only and audio-visual replies are paraphrases). If you only remember one eval lesson, remember those two.
How Griffin is put together
Tavus describes a two-part system.
1. Continuous Conversational Modeling. Incoming audio and video are reassessed on sub-second mini-turns. At each tick the engine decides whether to stay quiet, nod, backchannel, take the floor, or yield. Outputs are not only words. They are controls for stance, emotion, face, gaze, and gesture. A pause for thought is not supposed to be treated as end-of-turn. Perception keeps running while the PAL speaks.
2. Audio-visual generation. Streaming speech and streaming video follow those controls as they arrive.
Speech path, as published:
- Fast autoregressive diffusion transformer (VDiT) on an encoded speaker prefix (about 10 seconds of reference audio for a clone).
- Tavec: a convolutional autoencoder that maps 48 kHz audio to a continuous latent — 40 values per frame, 100 frames per second, no codebooks. Fully causal decoder. Packets as small as 10 ms. A minute of speech is 6,000 latent frames versus 2.88 million raw samples.
- Generation and playback overlap. A mid-utterance control can change the next frames, not the next reply.
Video path, as published:
- Previous Tavus renderers leaned on 3D priors and struggled with hands, large body motion, and dynamic backgrounds.
- Griffin generates every pixel from one reference image, including chair motion, shadows, and background.
- A large bidirectional diffusion teacher is distilled in three stages: Distribution Matching Distillation → teacher-forced autoregression → Self-Forcing on the model's own history so long rollouts drift less.
- 720p in 320 ms chunks (VAE 8× temporal compression: one latent → eight frames at 25 fps). Few-step generator, three diffusion steps per latent.
- True audio-to-video latency on H100s: 0.43 seconds average, which Tavus says is half the next published streaming-diffusion baseline. No lookahead audio.
The company also reports the isolated video generator first on DOVER, FID, and THEval, and second on LSE-C lip-sync at 7.27, because LSE-C rewards exaggerated mouth motion.
Capability demos in the post (open-ended sessions with people who had never used Griffin): overlapping celebration after a promotion, Simon Says that copies a gesture only when the words match, co-constructed stories with barge-in, a Rubik's cube coach that watches the cube, soldering guidance that speaks when the next step is due rather than at every silence. Timings on the marketing diagrams are illustrative, which Tavus labels in the caption.
What this changes for what you build
If you already have Tavus PALs: do not rewrite the stack this week. Phoenix, Raven, and Sparrow are still the production path. Griffin-Lite is a wait-list research surface. Plan for a later swap of the interaction model, not a Monday migration.
If you are choosing a face + voice architecture: write down which job you need.
| Job | What to use in October 2026 | Why |
|---|---|---|
| Full-duplex voice in your own app | GPT-Live-1 in the API | Audio HIM-adjacent; no Tavus video loop |
| Face + overlays you can fork | HeyGen LiveAvatar demos | Cascade you can run; MIT starter |
| Camera agents on a line, not a face | NVIDIA VSS 3.3 vision-agent compose | Perception for incidents, not a conversational HIM |
| Face-to-face research preview | Griffin-Lite if you are a trusted tester | Not a customer SKU |
If you sell or staff video conversations: this is the jobs post meeting the interface. Outbound spray was already dying on text. A HIM that can hold a one-minute call without being clocked raises the floor for AI SDR / interview / L&D surfaces Tavus already markets — after disclosure is solved. Until then, a "human-like PAL" in a regulated workflow is a liability, not a demo win.
If you teach or hire for judgment: a face that nods on time is a new way to skip the hard part. Rehearsing a difficult conversation with a model that reacts like a person is useful if the human still does the live conversation. It is a meat-proxy machine if the rehearsal replaces the meeting and nobody checks the facts.
If you handle money or access on a call: keep the habit from the Hong Kong deepfake wire case. Out-of-band verify. A 48% "felt human" rate on a one-minute protocol is a warning about channel trust, not a reason to relax it. Tavus delayed GA for that reason. California already banned workplace emotion-inference AI in a different lane — HR tools that read faces. A HIM that performs emotion is the product-side twin of that policy fight.
If you write safety claims: the same week as this launch, the FTC confirmed a product-risk probe of major labs. Griffin's own post is the rare vendor note that withholds the SKU because it works too well at looking human. If you ship any avatar, match the marketing sentence to a disclosure you can show.
Honest limitations (including Tavus's)
- Not GA. Trusted testers. Form, not an API changelog.
- Short Turing protocol. One minute, one prompt, company study. Flow was the weakest rating.
- Marketing vs NVIDIA. "37% ahead" and "12x closer" are derived. Quote 3.83 / 3.73 / TOR.
- Perception comparisons mix configs. Tavus flags MiniCPM-o 4.5 and gpt-realtime as audio-only on the cited perception numbers.
- Latency numbers are H100, audio-to-video, for the generator in isolation versus published streaming diffusion models — not a promise of end-to-end call latency on your laptop.
- Identity and voice clone from a photo plus ~10 s of audio is a capability and a misuse surface.
- Existing customers stay on the cascade-plus-specialists stack. Phoenix / Raven / Sparrow.
Tavus credits Baseten, Daily, and Cerebrium for preview infrastructure, Queen Mary University of London for research collaboration, and NVIDIA for VideoFDB. Citation they request: Tavus Research, "Griffin: The First Human Interaction Model", Tavus, October 2026.
What to do this week
- If you are not a trusted tester, do not block a roadmap on Griffin. Keep the PAL on Phoenix / Raven / Sparrow.
- If you need duplex voice now, ship GPT-Live-1 and add a face later.
- If you need a face plus UI cards, fork the LiveAvatar + HyperFrames starter.
- Add a disclosure line to any avatar you already run. Tavus is designing "safe disclosure" because the study showed people did not think to ask.
- For any call that can move money, out-of-band confirm — same rule as the $25.6M video-call fraud write-up.
- If you evaluate vendors, ask for VideoFDB track + TOR + n, not a percentage-ahead slide.
Related on explainx.ai
- Are we all meat proxies now? — unread relay when the interface feels like a person
- AI is taking us to a point of no return — skill you stop practicing
- The jobs AI is killing: outbound is dead — where a face-to-face SDR lands
- HeyGen LiveAvatar + GPT-Live-1 demos — cascade you can run today
- GPT-Live-1 in the API — audio full-duplex without Griffin
- Deepfake fraud: $25.6 million video call — why deception delay is the story
- FTC probe of OpenAI and Anthropic — product-risk weather around a HIM launch
- California emotion-tracking ban — policy on faces at work
Official: Griffin research post (Tavus, Oct 1, 2026) · NVIDIA VideoFDB · VideoFDB paper, arXiv 2605.30256 · Launch thread
News post reflecting Tavus's October 1, 2026 Griffin announcement and NVIDIA VideoFDB leaderboard figures as of October 2, 2026. Face-to-face counts (26/54, 1/41), VideoFDB scores, Tavec/latency/chunk sizes, and the testers-only restriction match those primary write-ups. Marketing percentages are explained, not used as scores. Specifications and access will change — verify against Tavus and NVIDIA before you plan a build.
