MUR: Multimodal Unified Refinement for Multimedia Recommendation
摘要
In this paper, we propose a novel Multimodal Unified Refinement for Multimedia Recommendation (MUR) framework to address three critical challenges in multimodal recommendation: multimodal domain shift, preference-irrelevant multimodal noise, and incomplete multimodal fusion. The MUR framework incorporates three key components to enhance multimodal features for improved recommendation performance. First, a Multimodal Contrast Layer aligns features with recommendation-specific distributions to mitigate the multimodal domain shift between the pre-trained model and the recommendation task. Second, a Modal Purification Layer focuses on removing impurity noise, such as background elements in images or unrelated text in titles, while preserving key features. Finally, a Dual-View Multimodal Fusion Layer ensures a comprehensive understanding of user preferences and item details by pooling diverse insights from multiple modalities. Through extensive experiments on three public datasets, we demonstrate the effectiveness of the MUR framework to tackle the challenges faced by multimodal recommendation in comparison to existing methods, and pave the way for more accurate and robust recommendations in e-commerce and other domains.