错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-grained Cross-Modal Feature Fusion Network for Diagnosis Prediction

  • Ying An,
  • Zhenrui Zhao,
  • Xianlai Chen

摘要

Electronic Health Record (EHR) contains a wealth of data from multiple modalities. Utilizing these data to comprehensively reflect changes in patients’ conditions and accurately predict their diseases is an important research issue in the medical field. However, most fusion approaches employed in existing multimodal learning studies are excessively simplistic and often neglect the hierarchical nature of intermodal interactions. In this paper, we propose a novel multi-grained cross-modal feature fusion network. In this model, we first use hierarchical encoders to learn multilevel representations of multimodal data and a specially designed attention mechanism to explore hierarchical relationships within a single modality. Afterward, we construct a fine-grained cross-modal clinical semantic relationship graph between code and sentence representations. Then we employ Graph Convolutional Networks (GCN) on this graph to achieve fine-grained feature fusion. Finally, we use attention mechanisms to fully learn the contextual interactions between visit-level multimodal representations, and realize coarse-grained feature fusion. We evaluate our model on two real-world clinical datasets, and the experimental results validate the effectiveness of our model.