<p>Medical report generation has been a challenging task due to information bias between image and text, difficulty in locating lesion regions, and long-tailed data distribution. To address these, we propose a knowledge-enhanced cross-modal alignment network (KECAN). It consists of three modules, visual feature extraction (VFE), knowledge enhanced network (KEN) based on ophthalmic ultrasound knowledge graph (OU-KG), and cross-modal alignment network (CAN). KEN designs the multi-attention fusion mechanism including the knowledge graph attention (KGA) and knowledge-visual attention (KVA), which enhance visual information with prior knowledge to better discover the key abnormal regions in image. CAN is presented to align the bias between the small lesion in image and less text descriptions. The cross-entropy loss based on term frequency-inverse document frequency (TF-IDF) is proposed to improve the long tailed data learning. The proposed method improves the generation quality and the identification accuracy of abnormal due to the highlight association knowledge and fine-grained alignment. Experiment results on large ophthalmic dataset show that the proposed method achieves better performance compared with state-of-the-art methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

KECAN: knowledge-enhanced cross-modal alignment network for ophthalmic report generation

  • Jing Wang,
  • Mingyu Shi,
  • Junyan Fan,
  • Yanzhu Zhang,
  • Ruiping Wang

摘要

Medical report generation has been a challenging task due to information bias between image and text, difficulty in locating lesion regions, and long-tailed data distribution. To address these, we propose a knowledge-enhanced cross-modal alignment network (KECAN). It consists of three modules, visual feature extraction (VFE), knowledge enhanced network (KEN) based on ophthalmic ultrasound knowledge graph (OU-KG), and cross-modal alignment network (CAN). KEN designs the multi-attention fusion mechanism including the knowledge graph attention (KGA) and knowledge-visual attention (KVA), which enhance visual information with prior knowledge to better discover the key abnormal regions in image. CAN is presented to align the bias between the small lesion in image and less text descriptions. The cross-entropy loss based on term frequency-inverse document frequency (TF-IDF) is proposed to improve the long tailed data learning. The proposed method improves the generation quality and the identification accuracy of abnormal due to the highlight association knowledge and fine-grained alignment. Experiment results on large ophthalmic dataset show that the proposed method achieves better performance compared with state-of-the-art methods.