Stronger interaction brings better performance: fine-grained alignment between different modalities for MNER
摘要
Multi-modal named entity recognition (MNER) aims to enhance entity representations and boost recognition performance by leveraging visual knowledge from images. However, previous methods mainly faced two limitations: insufficient fine-grained alignment between text and image regions, and difficulty in filtering irrelevant visual information. To address these issues, we propose fine-grained alignment (FGA) between different modalities for MNER. Firstly, FGA extracts visual objects and calculates accurate fine-grained correlations with textual entities. Secondly, we introduce a contrastive learning strategy to integrate relevant visual knowledge, thereby mitigating noise from irrelevant image regions. Finally, we propose a dynamic weight optimization mechanism to adaptively enhance or filter multi-modal information based on its relevance. Experimental results on three benchmark datasets show that our method achieves competitive performance compared to existing methods, validating the benefits of precise and adaptive multi-modal information fusion.