chiprook
← AI
AIOctober 1, 2026, 22:54

UniEvo-VL: multimodal model learns from its own image critiques

A new paper introduces UniEvo-VL, an on-policy self-distillation method that turns image critiques into training supervision: the student matches an EMA teacher conditioned on a critique-enriched prompt while seeing only the original one. Built on Qwen-Image-2512, it lifts GenEval from 0.747 to 0.808 and GenEval2 Soft-TIFA from 32.97 to 35.53.

UniEvo-VL: multimodal model learns from its own image critiques
#Qwen
Read next
AI

TaichuAI Open-Sources ZDTaichu5.0-9B Spatial Multimodal Model

AI

Alibaba launches Qwen Image 2.1 with just 7B parameters

AI

Meituan ships LongCat-2.5-Preview: 1.6T MoE multimodal agent with 1M context

AI

Shanghai AI Lab releases Intern-S2-397B science multimodal model