Multimodal recommender systems, which integrate heterogeneous features (e.g., visual, textual) with user interaction patterns, have become a pivotal research direction in recommendation tasks. While recent methods leverage auxiliary graphs constructed from multimodal content to capture item correlations and combine them with collaborative filtering, they suffer from critical limitations: (1) Overreliance on modality-similarity-based item-item relationships introduces noise and restricts the learning of robust, comprehensive item representations; (2) Naively applying behavioral data to refine multimodal features ignores the inherent sparsity and bias in user behavior signals. To address these challenges, we propose TEMOO (Two-stagE Multi-view cOntrastive fusiOn), a framework that learns denoised and generalized item correlations through two key innovations. First, a multi-view information encoder constructs three distinct feature perspectives: (a) modality-driven item-item relations, (b) behavior-driven item-item relations, and (c) a purified user-item interaction view. Second, a two-stage contrastive fusion mechanism is designed to jointly model shared and complementary information across views: (1) dual alignment of modality and behavior-based item relations to suppress multimodal noise, and (2) adaptive fusion of denoised collaborative signals to filter noisy edges in behavior-derived graphs. Extensive experiments on four real-world datasets demonstrate that TEMOO achieves superior performance over state-of-the-art methods in accuracy, training efficiency, and memory efficiency.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Denoising Multimodal Recommendation with Two-Stage Multi-view Contrastive Fusion

  • Haibo Liu,
  • Yunlong Zhou,
  • Limin Wu,
  • Yuchen Zhuang,
  • Jinglian Liu

摘要

Multimodal recommender systems, which integrate heterogeneous features (e.g., visual, textual) with user interaction patterns, have become a pivotal research direction in recommendation tasks. While recent methods leverage auxiliary graphs constructed from multimodal content to capture item correlations and combine them with collaborative filtering, they suffer from critical limitations: (1) Overreliance on modality-similarity-based item-item relationships introduces noise and restricts the learning of robust, comprehensive item representations; (2) Naively applying behavioral data to refine multimodal features ignores the inherent sparsity and bias in user behavior signals. To address these challenges, we propose TEMOO (Two-stagE Multi-view cOntrastive fusiOn), a framework that learns denoised and generalized item correlations through two key innovations. First, a multi-view information encoder constructs three distinct feature perspectives: (a) modality-driven item-item relations, (b) behavior-driven item-item relations, and (c) a purified user-item interaction view. Second, a two-stage contrastive fusion mechanism is designed to jointly model shared and complementary information across views: (1) dual alignment of modality and behavior-based item relations to suppress multimodal noise, and (2) adaptive fusion of denoised collaborative signals to filter noisy edges in behavior-derived graphs. Extensive experiments on four real-world datasets demonstrate that TEMOO achieves superior performance over state-of-the-art methods in accuracy, training efficiency, and memory efficiency.