<p>Automatically generating diagnostic reports from medical imaging scans provides substantial clinical benefits by improving patient experiences and reducing diagnostic pressures on healthcare professionals. Current medical report generation models primarily focus on clear scans and typically produce template sentences that convey limited informative content. However, in real-world medical contexts, scans often exhibit a lack of clarity, and reports adhere to a specific terminological and logical structure, which renders existing methods insufficient for effectively describing areas with abnormalities. To address these challenges, this paper introduces a novel model for automatic medical report generation based on fine-grained semantic alignment and cross-modal enhancement, referred to as FVA-CD. The model achieves cross-modal semantic alignment between visual and textual modalities through multimodal pre-training, thereby enhancing its capability to accurately model detailed regions of interest. Additionally, by integrating the generated medical terms into a cross-modal attention module, the encoder improves its ability to recognize and interpret abnormal semantics. Furthermore, to facilitate the generation of abnormal terminology, a conditional decoder is designed to leverage both abnormal terms and the overall report structure, enabling the dynamic generation of target words. Experimental results from the MIMIC-CXR and ROCO public datasets demonstrate significant improvements across multiple descriptive metrics with the proposed model.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

From coarse to grain: automated medical report generation based on fine-grained semantic alignment and cross-modal enhancement

  • Yuhao Tang,
  • Fei Tao

摘要

Automatically generating diagnostic reports from medical imaging scans provides substantial clinical benefits by improving patient experiences and reducing diagnostic pressures on healthcare professionals. Current medical report generation models primarily focus on clear scans and typically produce template sentences that convey limited informative content. However, in real-world medical contexts, scans often exhibit a lack of clarity, and reports adhere to a specific terminological and logical structure, which renders existing methods insufficient for effectively describing areas with abnormalities. To address these challenges, this paper introduces a novel model for automatic medical report generation based on fine-grained semantic alignment and cross-modal enhancement, referred to as FVA-CD. The model achieves cross-modal semantic alignment between visual and textual modalities through multimodal pre-training, thereby enhancing its capability to accurately model detailed regions of interest. Additionally, by integrating the generated medical terms into a cross-modal attention module, the encoder improves its ability to recognize and interpret abnormal semantics. Furthermore, to facilitate the generation of abnormal terminology, a conditional decoder is designed to leverage both abnormal terms and the overall report structure, enabling the dynamic generation of target words. Experimental results from the MIMIC-CXR and ROCO public datasets demonstrate significant improvements across multiple descriptive metrics with the proposed model.