Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning
Yijia Fan, Ziqi Huang, Zhongang Cai, Yan Li, Zimo Wen, Wanqi Yin, Haiwen Diao, Ziwei Liu
- Published
- Sep 28, 2026 — 17:59 UTC
Problem
Unified multimodal models require the capability to jointly learn reflection text and image generation to facilitate effective self-repair. This paper addresses this gap by proposing a novel approach that integrates reinforcement learning into the training of these models. The work is presented as a preprint and has not undergone peer review.
Method
The authors introduce the UMM-Reflection model, which employs reinforcement learning (RL) to optimize the completion of reflection trajectories. The model utilizes the BAGEL dataset as its data source. The training mechanism involves sibling trajectories that share a common initial image, allowing for a comparative analysis of reflection strategies through group-relative advantage. This approach enables trajectory-level advantage updates, which refine reflection tokens and facilitate flow-based revisions. The architecture is designed to enhance the model's ability to generate coherent and contextually relevant reflections across modalities.
Results
The proposed method demonstrates significant improvements over the standard fine-tuning (SFT) baseline across multiple benchmarks:
- GenEval: 12.05 points improvement vs SFT
- WISE: 10.97 points improvement vs SFT
- OneIG-Bench: 3.48 points improvement vs SFT
- T2I-CompBench++: 4.63 points improvement vs SFT
These results indicate the effectiveness of the UMM-Reflection model in enhancing multimodal reflection generation capabilities.
Limitations
The authors do not report any limitations in their work. However, as a preprint, the lack of peer review may imply potential oversights or unaddressed issues that could be identified in a formal review process.
Why it matters
The implications of this research are significant for downstream applications in multimodal AI systems, particularly in enhancing the self-repair capabilities of models that generate and reflect on both text and images. By improving the integration of reflection mechanisms, this work paves the way for more robust and contextually aware AI systems, which can lead to advancements in areas such as interactive AI, content generation, and automated reasoning.
By Callan Zhang · Sep 28, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: arXiv cs.AI
