Enhancing Chinese Multimodal Entity Linking with CLIP-RoBERTa and Contrastive Learning
摘要
Multimodal Entity Linking (MEL) is a vital task in natural language processing, aiming to associate ambiguous mentions in multimodal data—such as text and images with entities in a knowledge base (KB). However, existing methods focus primarily on English corpora, limiting their applicability to Chinese data. This challenge is further exacerbated by the scarcity of Chinese text-image datasets and semantic inconsistencies between English and Chinese languages, making it difficult to adapt English-based models for Chinese contexts. Short text labels often lack semantic richness, while noise in visual data introduces irrelevant features, reducing linking accuracy. To address these challenges, we propose MCR, a contrastive learning-based Multimodal Entity Linking method tailored for Chinese datasets. MCR leverages CLIP-RoBERTa for deep feature learning and incorporates contrastive learning to strengthen feature relationships, enabling more accurate linking of multimodal data. Additionally, we contribute a high-quality Chinese MEL dataset focused on ethnic minority elements, addressing a critical resource gap. Experimental results on this new dataset and other benchmarks demonstrate that MCR outperforms existing methods across multiple metrics, establishing it as a robust solution for MEL in both Chinese and English domains.