D-DPDG: Diffusion-based dual-graph attention with dual-path feature extraction for multimodal recommendation
摘要
In recent years, multimodal recommender systems have gained significant attention as a promising solution to address complex recommendation scenarios. However, a major challenge lies in effectively capturing the intricate relationships in user-item interactions. Since multimodal data is derived from various sources, it often contains noise, which can negatively impact the model’s ability to accurately understand user preferences. Traditional methods typically aggregate information from neighboring nodes in the graph structure in a static manner, which fails to consider the varying importance of these nodes, particularly when there is heterogeneity among them. To overcome these challenges, this paper develops a multimodal recommendation system that leverages dual-graph attention and dual-path feature extraction for comparative learning. The proposed method enhances the feature information of each modality by applying multi-level feature augmentation, reduces noise in the user-item interaction graph using a diffusion model, and dynamically captures the importance of different nodes through dual-graph attention fusion. Experimental results demonstrate that the proposed method outperforms existing benchmarks on four datasets: TikTok, Amazon-Baby, Amazon-Sports, and Amazon-Clothing.