Automated chest X-ray diagnosis often relies on unimodal models, such as ResNet and CXR-BERT, which struggle to integrate heterogeneous imaging and textual features. This limitation reduces robustness, particularly in cases of pathological feature overlap and anatomical variations. To address these challenges, we propose a multimodal classification framework that combines Lie Group feature learning with multimodal dynamic attention. First, we introduce a Lie Group feature extractor (LGFE) module that maps local gradient and texture features onto Lie Group manifolds, ensuring geometrically invariant representations. Second, we develop a multimodal dynamic attention mechanism (MDAM) to align imaging regions with textual keywords using Lie Group distance metrics, enhancing semantic correlation. Finally, we design a Lie Fisher classifier to improve intra-class cohesion and inter-class separation in the manifold space. Experiments on the MIMIC-CXR dataset demonstrate that our model achieves a prediction accuracy of 84.0% and an average AUC score of 93.5% across 14 distinct impression categories, outperforming baseline models by 3.9% and 7.3%, respectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Lie Group-Based Multimodal Dynamic Attention for Chest X-ray Diagnosis

  • Tianchen Fang,
  • Guiru Liu

摘要

Automated chest X-ray diagnosis often relies on unimodal models, such as ResNet and CXR-BERT, which struggle to integrate heterogeneous imaging and textual features. This limitation reduces robustness, particularly in cases of pathological feature overlap and anatomical variations. To address these challenges, we propose a multimodal classification framework that combines Lie Group feature learning with multimodal dynamic attention. First, we introduce a Lie Group feature extractor (LGFE) module that maps local gradient and texture features onto Lie Group manifolds, ensuring geometrically invariant representations. Second, we develop a multimodal dynamic attention mechanism (MDAM) to align imaging regions with textual keywords using Lie Group distance metrics, enhancing semantic correlation. Finally, we design a Lie Fisher classifier to improve intra-class cohesion and inter-class separation in the manifold space. Experiments on the MIMIC-CXR dataset demonstrate that our model achieves a prediction accuracy of 84.0% and an average AUC score of 93.5% across 14 distinct impression categories, outperforming baseline models by 3.9% and 7.3%, respectively.