错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Semantic and Visual Attention-Driven Multi-LSTM Network for Automated Clinical Report Generation

  • Cheng Huang,
  • Junhao Shen,
  • Beichen Hu,
  • Mohammad Ausaf Ali Haqqani,
  • Tsengdar Lee,
  • Karanjit Kooner,
  • Ning Zhang,
  • Jia Zhang

摘要

Medical image processing has gained significant momentum in recent years. Latest advancements in machine learning and deep learning have enabled the AI-powered generation of medical image reports. However, some limitations remain. For example, generated reports may be lengthy without highlighting anomalies as desired; some minor features might be neglected, which fails in fine-grained labeling.To tackle the aforementioned challenges, this paper presents a Semantic and Visual Attention-Driven Multi-LSTM Network (SVAML), a novel framework tailored to enhance medical image report generation. In particular, SVAML introduces a Double-Weighted Multi-Head Attention mechanism with a crafted weight function to learn patterns of how to focus on describing important impressions from medical images. In addition, SVAML devises a Label Discriminator (LD), a module to learn intricate features to support more sensitive multi-label classification. Extensive experiments over two known public datasets, the IU X-ray dataset and the PEIR Gross dataset, have demonstrated the effectiveness of the presented SVAML framework.