Prism brings dynamic sparse attention to joint video-audio generation
Prism is a dynamic sparse-attention framework for training high-resolution joint video-and-audio generation models. The authors report 2.5x faster training than full attention with higher generation quality; 720p inference needs one 80 GB GPU, while 1080p and 2K need four or more.
- Prism splits video into spatiotemporal zones and adapts attention blocks to local content
- Authors report 2.5x faster training than full attention with higher generation quality
- 720p inference requires one 80 GB GPU; 1080p and 2K need four or more
- Native training needs at least 32 80 GB GPUs for 720p and 64 for 1080p/2K
Read next
AI