Look, Imitate and Refine: A Hierarchical Multimodel Retrieval Augmented Vision-Language Model for Radiology Report Generation
摘要
The task of Radiology Report Generation (RRG) aims to automatically produce diagnostic reports that accurately describe findings in radiological images, assisting clinicians in decision-making and alleviating the workload of radiologists. Existing approaches primarily use generative models that rely on visual features extracted from the images to generate text. However, these methods face significant challenges in aligning the generated text with the visual content and maintaining coherence across complex descriptions. To address these limitations, we propose a novel multimodel retrieval-augmented framework, Look, Imitate, and Refine, for radiology report generation. Our approach is divided into three stages: Look—enhances visual understanding through image retrieval and feature extraction; Imitate—utilizes retrieved reports to guide the generation of a draft report, ensuring contextual relevance; and Refine iteratively refines the draft report by leveraging additional textual context, enhancing semantic alignment and overall quality. Extensive experiments on the IU-XRay and MIMIC-CXR datasets demonstrate that our framework outperforms state-of-the-art methods, significantly improving report quality and diagnostic accuracy.