REVERIE+: Generalized Reflective Instruction Tuning for Hallucination Mitigation in Advanced VLMs
摘要
Large Vision-Language Models (LVLMs) demonstrate strong capabilities but remain susceptible to hallucinations. To address this limitation, we propose reflective instruction tuning, which explicitly trains models to reflect. Rather than producing only a final answer, the model is supervised to generate reflective rationales that justify the correct prediction and explain why plausible alternatives are incorrect. To remain effective for advanced LVLMs deployed in diverse real-world scenarios, reflective supervision must be more fine-grained and broader in domains. We therefore propose REVERIE+ (ReflEctiVE RatIonalE), an extension of REVERIE substantially expanded in domain diversity, task complexity, and annotation richness, tailored to advanced LVLMs. Built on the R1-Onevision data foundation, REVERIE+ broadens domain coverage and increases task difficulty, while improving annotation reliability by (i) leveraging multiple models to mine and label diverse negative answers, and (ii) employing a strong commercial LVLM to generate higher-quality reflective rationales. These changes provide richer negative supervision and more accurate reflection signals, making REVERIE+ better suited for hallucination mitigation in more capable models. Extensive experiments across representative hallucination and general multimodal benchmarks show consistent performance gains, validating that the enriched negative supervision and precise reflection signals of REVERIE+ offer a scalable and effective approach for mitigating hallucinations in advanced LVLMs.