Similarity Retrieval and Medical Cross-Modal Attention Based Medical Report Generation
摘要
Medical report generation is a time-consuming and knowledge-intensive task performed by radiologists to describe various regions within medical images. Writing report manually is prone to subjective bias and errors. Consequently, medical report generation automatically has become an important research direction in the field of artificial intelligence. While recent report generation methods have achieved relatively fluent medical reports, several challenges remain: (1) They overlook valuable semantic information from similar cases, which results in insufficient information for accurate reporting; (2) Deep exploration of medical features is lacking, which hampers the understanding of medical terminology and compromises disease prediction accuracy. To address the aforementioned issues, this paper proposes a Similarity Retrieval and Medical Cross-modal Attention based Medical Report Generation Network (SRMCAN). By employing content-based similarity retrieval, SRMCAN filters out interfering information in relevant semantic features, which serves as a complementary feature for the model. SRMCAN constructs a fine-grained alignment loss function, taking similar cases as hard negative samples to enhance the dynamic interaction between cross-modal disease features. A Medical Cross-modal Attention mechanism is designed to capture second-order interactions between cross-modal features, incorporating coordinate attention to calculate attention distributions in two spatial directions. The Medical Cross-modal Attention strengthens the model’s understanding and reasoning ability for medical information, improving the accuracy and professionalism of the generated reports. Experimental results on the IU X-Ray dataset demonstrate that SRMCAN improves the fluency, accuracy, and professionalism of medical reports, providing radiologists with more valuable reference medical reports.