Kandinsky 6.0 Video is an open, MIT-licensed pair of video models that generate picture and sound in one pass. Kandinsky Lab published the code, checkpoints, and an arXiv technical report on October 6, 2026. There are two sizes: a 29-billion-parameter Pro for cinematic quality and a 3-billion-parameter Lite for quick iteration. Each makes five-second clips with synchronized 44 kHz audio, and a separate super-resolution stage raises the picture to 1920 by 1080.
That combination matters because open video models have mostly been silent. Adding speech, effects, and lip-sync usually means a second model and a manual alignment step. Kandinsky 6.0 trains both streams together. The headline for builders is the license: MIT, with no territory carve-outs. That is a sharp contrast to MiniMax H3, whose open weights exclude the US, EU, UK, and South Korea.

Quick answers
| Question | Short answer |
|---|---|
| What was released? | Kandinsky 6.0 Video Pro (29B) and Lite (3B), plus a video super-resolution model |
| License? | MIT for code, checkpoints, and diffusers integration, per the repository and report |
| What does it generate? | Five-second clips with synchronized 44 kHz audio, text-to-audio-video and image-to-audio-video |
| Output resolution? | Up to Full-HD (1920x1080) via built-in super-resolution |
| What hardware? | An NVIDIA GPU and Python 3.13 or 3.14; Pro HD runs 292 seconds on an H100 (non-distilled) |
| Can I try it hosted? | Yes: a Hugging Face Space demo; also vLLM-Omni and ComfyUI nodes |
| Biggest caveat? | The quality comparison is the authors' own human study; clips are only five seconds |
What exactly did Kandinsky Lab release?
According to the project repository, Kandinsky 6.0 Video consists of two diffusion models. Both support text-to-audio-video (T2AV) and image-to-audio-video (I2AV). In the second mode you hand the model a still image and a prompt, and it animates the image and generates a matching soundtrack.
The Hugging Face paper page for the report lists the team as a large collective, with 88 contributors on the arXiv entry. The repository also links a diffusers integration, so the models can be loaded through the same Python library many readers already use for image models. For the technical background on the earlier generation, see the Kandinsky 5.0 report and the diffusers pipeline documentation.
There is also a dedicated video super-resolution repository. That split is practical: the base models generate at a lower resolution, and the upscaler is a separate plugin you can swap or skip.
How does it make audio and video together?
The report describes a dual-stream CrossDiT architecture. One stream is a pretrained video diffusion transformer. The second is a newly trained audio stream. The two are connected by bidirectional cross-attention, meaning video tokens can look at audio tokens and the reverse at every layer where they are linked. That is how the model is supposed to keep a closing door and its thud on the same frame, or lip movement aligned with speech.
The training recipe, as the authors summarize it, has several stages:
- Pretrain the audio stream on large-scale audio corpora.
- Train the video and audio streams jointly on paired audio-video data.
- Supervised fine-tuning on curated examples.
- Reinforcement-learning post-training.
- Distillation down to 10 sampling steps for a faster variant.
Two numbers stand out. The RL stage reportedly cut speech word error rate by 47 percent, which is a direct measure of how intelligible generated speech is when transcribed. The report also lists a 10-step distillation stage, and the quick-start downloads the pro-distill checkpoint by default.
For a plain-language refresher on how diffusion transformers differ from autoregressive models, our guide to video generation tools across Sora, Runway, and Kling covers the basics of the category.
How good is it, really?
The authors report a human evaluation in which Kandinsky 6.0 Pro clearly beats Kandinsky 5.0 Pro and is preferred over LTX 2.5 on visual quality, motion realism, visual prompt following, artifact reduction, and overall task solving. On speech quality the report claims a statistically significant advantage over LTX 2.5.
Read that carefully. This is a first-party study run by the team that built the model. It does not compare against the strongest closed systems, and the report describes the Pro model as "competitive" with leading models on speech rather than ahead of all of them. We have not independently reproduced the evaluation, and independent leaderboards typically take days to weeks to list a new open release. Wait for those before treating the ranking as settled.
What the claim does tell you: for an openly licensed model you can run yourself, the team is aiming at the same tier as current open competitors like LTX 2.5, with audio treated as a first-class output instead of an add-on.
What do you need to run it?
The repository is explicit about the basics. You need an NVIDIA GPU and Python 3.13 or 3.14, plus uv and just. The install flow uses the just task runner:
git clone https://github.com/kandinskylab/kandinsky-6
cd kandinsky-6
just setup
just download pro-distill
just generate "A street musician plays violin in a rainy alley, close-up, ambient rain and reverb"
Weights are cached under ~/.cache/kandinsky, and each run writes a dated output folder containing the video, the expanded prompt, logs, and the configuration, which makes runs easy to reproduce and compare. Prompts are automatically expanded before generation, so keep the original short and descriptive.
The per-GPU timings in the README are worth reading before you rent hardware. For the non-distilled Pro model at HD, the README lists 292 seconds on an H100 and 2,716 seconds on an RTX 5060 Ti for a 5-second clip, excluding weight loading and encoding. The attention backend also varies by card: Hopper GPUs use FlashAttention 3, while Ampere, Ada, and consumer Blackwell cards need SageAttention 2.2.0, which you compile yourself.
| Choice | Best for | Trade-off |
|---|---|---|
| Pro (29B), distilled | Highest quality, final renders | Slow on consumer cards, heavy memory use |
| Lite (3B) | Fast iteration, prompt testing | Lower fidelity than Pro |
| Hugging Face Space | No GPU, quick tests | Shared demo, Pro distill only, data leaves your machine |
| ComfyUI node package | Existing node-based workflows | Separate node installs via ComfyUI Manager |
| vLLM-Omni | Serving at scale | Requires an inference engineering setup |
For a first experiment, we would start with Lite or the hosted demo to learn how prompts behave, then move to Pro only for clips you intend to keep.
Why the MIT license is the real story
Open video models come with very different terms. Some restrict regions, some forbid using outputs to train other models, some require a commercial agreement above a revenue threshold. Our explainer on choosing open-weight versus closed models walks through why license text matters as much as benchmark scores when you plan a product.
MIT is about as permissive as it gets: use, modify, and redistribute, including commercially, with attribution and the license notice preserved. Practically, that means a studio, a startup, or a hobbyist can fine-tune Kandinsky on their own footage and ship the result. That is a different proposition from the gated terms we flagged in MiniMax H3 and from the hosted-only access of Seedance 2.5.
One honest caution: an MIT license on the weights does not remove other obligations. If your outputs feature real people's voices or likenesses, disclosure and consent rules still apply, and some jurisdictions are adding synthetic-media labeling laws, as we covered in the New York AI video disclosure law post.
Where does it sit among open and closed video models?
Video generation has moved quickly in 2026. Alibaba's Wan 3.0 and HappyHorse 1.1 pushed Chinese video models forward. MiniMax's fast H3 variant showed faster-than-playback generation on Blackwell hardware. On the closed side, Google's Gemini Omni Flash and ByteDance's Seedance 2.5 offer longer clips and higher peak quality through paid APIs.
Against that field, Kandinsky 6.0 has a clear niche:
- Audio and video in one model, rather than a silent clip plus a separate audio tool.
- Permissive license with no territory restrictions.
- A usable small model, the 3B Lite, which is rare in a field where open releases are often 30B and above.
- Broad tooling on day one: diffusers, vLLM-Omni, ComfyUI, and a public demo.
It also has clear limits. Five seconds is short next to closed models advertising clips up to 30 seconds. Full-HD comes through an upscaling stage, not native generation. And the Pro model is expensive in GPU time unless you use the distilled checkpoint or an H100-class card.
What people will ask next
Can it make a whole scene or a short film? Not directly. Five-second clips are building blocks. Teams stitch them with an editing workflow, use image-to-video to keep a character consistent between shots, and rely on an editor for pacing. Our guides to agentic video production show how multi-shot pipelines are assembled around short-clip models.
Is the audio good enough for dialogue? The report highlights lip-sync and a speech word error rate reduction from RL post-training, and claims a significant speech-quality edge over LTX 2.5. Quality in your language, accent, and noise conditions is something to test yourself. Start with short lines of dialogue and transcribe the output with a speech-to-text tool to measure intelligibility.
Will it run on my gaming GPU? Probably, slowly. The README lists benchmarks for seven GPU models across SD, HD, and Full-HD, and consumer cards appear in the table. A card with limited memory will push you toward Lite.
Is it safe to deploy? Open video-and-voice models raise obvious misuse risks, including impersonation. If you build on it, add disclosure labels, avoid cloning real voices without consent, and log generations. See our report on deepfake video-call fraud for what is at stake.
A practical first-hour plan
- Open the Hugging Face Space demo linked from the repository and try three prompts: a silent-looking scene, a scene with obvious sound, and a short spoken line.
- If the results fit your use case, run Lite locally with
just generateand compare it against the demo output. - Run the same prompt on the distilled Pro checkpoint and note the render time on your hardware.
- Use image-to-video with one of your own reference images to test consistency.
- Only then decide whether the super-resolution stage is worth the extra seconds for your delivery format.
Keep a log of prompts and seeds. The repository writes the expanded prompt and configuration for each run, which is enough to reproduce a clip later.
Bottom line
Kandinsky 6.0 Video is a credible, permissively licensed entry in the open video race, and its main differentiator is native audio generation in a model you can actually download. The quality ranking vs LTX 2.5 is the authors' own and should be confirmed by independent leaderboards. The five-second clip length and GPU cost are real constraints. If you build with video AI and care about a clean license, it belongs on your test list this week.
Related reading
- MiniMax H3: open video model with territory restrictions
- MiniMax fast H3 and real-time open video on Blackwell
- Alibaba Wan 3.0 video model
- Video generation AI complete guide: Sora, Runway, Kling
- How to choose open-weight vs closed AI models
- Seedance 2.5: 30-second 4K AI video
- Gemini Omni Flash video generation
Official sources: technical report on arXiv, Kandinsky 6 GitHub repository, Hugging Face paper page.
Details reflect the repository and report as of October 6, 2026. Check the repository for updated requirements, checkpoints, and license files.
