UniEvo-VL lets a multimodal model teach itself by critiquing its own images
UniEvo-VL: Self-Distillation Training for Multimodal Model Self-Improvement
UniEvo-VL is a self-evolving training recipe that turns a multimodal model's own critiques into privileged information: one model acts as both teacher and student, with the student seeing only the question and the teacher conditioned on the critique. Training minimizes the divergence between their denoising distributions along the student's own sampling trajectories. Built on Qwen-image-2512, it lifts GenEval from 0.747 to 0.808 and GenEval2 Soft-TIFA from 32.97 to 35.53, though gains are not uniform across tasks.
Instead of relying on a separate, often larger, teacher, we leverage their self-critiques as privileged information and ask a single multimodal model to act as both teacher and student with different contexts.