Enhancing Medical Image Report Generation Through Regional and Global Feature Interaction
摘要
Medical image report generation has gained significant attention as an essential task in automated diagnosis, aiming to reduce the workload of physicians and expedite the diagnostic process. However, existing methods predominantly rely on coarse-grained global features, neglecting the fine-grained regional analysis. This limitation potentially leads to the loss of crucial diagnostic information. To address this challenge, we propose the Regional-Global Interaction Network (RGI), an innovative architecture designed to provide more granular information to the model. The RGI introduces two key components: (1) a region selection agent that precisely locates and selects relevant image patches based on input regional word embeddings, and (2) a bidirectional cross-attention mechanism that fuses the selected regional features with global features. This approach enables the model to capture local details while retaining global context. Extensive experiments on the IU X-Ray dataset demonstrate that our RGI architecture significantly improves the accuracy of medical image analysis and report generation. The proposed method outperforms existing state-of-the-art models across multiple evaluation metrics, including BLEU and ROUGE scores.