Kandinsky 6.0 Video: open model generates video with synced audio
An arXiv pre-print published on October 4, 2026 introduced Kandinsky 6.0 Video, an open two-tier diffusion model that generates Full-HD video with synchronized 44 kHz audio. Weights are on Hugging Face and inference runs via Stability Cloud at about $0.03 per 5-second clip.
- Lite tier (~3B parameters) runs on a consumer GPU; Pro (~29B) fits a single A100-40GB
- Lip-sync MOS of 4.3 versus 3.6 for Sora-Lite, which lacks native audio
- A 5-second clip takes about 45 seconds on an RTX 4090 and costs roughly $0.02
- Weights, Dockerfile and a transformers-compatible pipeline are public on Hugging Face
Read next
AI