Artificial intelligence has transformed numerous fields, with a profound impact on medical imaging, a cornerstone of modern diagnostics. For radiologists, interpreting images is often time-consuming and complex, creating a need for AI-assisted solutions. However, the abstract nature of medical images and the cross-modal challenges between images and text have slowed progress in automated report generation. This paper presents the Dual Visual Feature-Driven Cross-modal Radiology Report Generation (DV-CRRG) network, which learns local lesion and grid features of medical images to align images with text through learnable queries, generating accurate reports. Experiments show that DV-CRRG achieves superior results on both the IU X-Ray English dataset and a proprietary Chinese CT dataset.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dual Visual Feature-Driven Cross-Modal Radiology Report Generation

  • Weili Zhang,
  • Fujiao Ju

摘要

Artificial intelligence has transformed numerous fields, with a profound impact on medical imaging, a cornerstone of modern diagnostics. For radiologists, interpreting images is often time-consuming and complex, creating a need for AI-assisted solutions. However, the abstract nature of medical images and the cross-modal challenges between images and text have slowed progress in automated report generation. This paper presents the Dual Visual Feature-Driven Cross-modal Radiology Report Generation (DV-CRRG) network, which learns local lesion and grid features of medical images to align images with text through learnable queries, generating accurate reports. Experiments show that DV-CRRG achieves superior results on both the IU X-Ray English dataset and a proprietary Chinese CT dataset.