错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multifactorial modality fusion network for multimodal recommendation

  • Yanke Chen,
  • Tianhao sun,
  • Yunhao Ma,
  • Huhai Zou

摘要

Multimodal recommendation systems aim to deliver precise and personalized recommendations by integrating diverse modalities such as text, images, and audio. Despite their potential, these systems often struggle with effective modality fusion strategies and comprehensive modeling of user preferences. To address these issues, we propose the Multifactorial Modality Fusion Network (MMFN). MMFN overcomes the limitations of previous models by following pivotal architectures. First, this novel approach employs three Graph Neural Networks (GNN) to extract foundational interactions and semantic information across modalities meticulously. Second, a Gated Multi-factor Semantic Sensor operates through a series of stacked gating units, guided by interaction embeddings, to extract features from modal embeddings deeply. Third, a User Preference-Oriented Modality Aligner, leveraging contrastive learning to synchronize user preferences with item features, thus enhancing the expressiveness of embeddings and the overall quality of recommendations. We demonstrate the marked superiority of MMFN in both performance and efficiency compared to traditional collaborative filtering methods and contemporary deep multimodal recommendation systems. Through comprehensive evaluations on the baby, sports, and clothing datasets, MMFN achieves significant gains in Recall@20 metrics, with improvements of 2.49%, 8.79%, and 24.51% over the following best baseline models. Additionally, MMFN also leads in training efficiency, outperforming most competing models. MMFN paves the way for future multimodal recommendation systems, leveraging the full spectrum of deep learning technologies.