UniEvo-VL: multimodal model learns from its own image critiques
A new paper introduces UniEvo-VL, an on-policy self-distillation method that turns image critiques into training supervision: the student matches an EMA teacher conditioned on a critique-enriched prompt while seeing only the original one. Built on Qwen-Image-2512, it lifts GenEval from 0.747 to 0.808 and GenEval2 Soft-TIFA from 32.97 to 35.53.
- GenEval rises from 0.747 to 0.808; GenEval2 Soft-TIFA from 32.97 to 35.53
- Built on Qwen-Image-2512 with a Qwen-VL feedback pipeline
- EMA teacher (decay 0.999) sees the revised prompt; student sees only the original
- Without verification, OCR drops from 0.771 to 0.761
Read next
AI