TaichuAI Open-Sources ZDTaichu5.0-9B Spatial Multimodal Model
TaichuAI released the ZDTaichu5.0-9B model on Hugging Face with ~9 billion parameters: Qwen3.5-9B language backbone and C-RADIOv4-H vision encoder, context up to 128K tokens, support for text, images, and video. The model targets spatial perception, embodied tasks, and agent work, under the NVIDIA Open Model License.
- ~9B parameters, context up to 128K tokens, input text, images, video
- Scores: ViewSpatial 62.50, RoboSpatial 56.00, TAU2-Bench 87.70, IFEval 93.70
- Weights on Hugging Face under NVIDIA Open Model License with Qwen3.5 Apache-2.0 notices
- For launch: custom vLLM 0.26.0 branch and Docker image with OpenAI-compatible API
Read next
AI