Online multi-modal sharing platforms have necessitated the integration of diverse data types such as visual, textual, and auditory to form a holistic user preference profile in personalized recommendation systems. Current multi-modal recommendation methods extract features from user-item interactions but often depend on custom tags and face high computational costs, particularly in visual feature processing, and struggle with effectively modeling long-range user-item dependencies. To overcome these challenges, we introduce a novel approach, Multi-perspective Attention Enhanced Multi-modal Self-supervised GANs for Recommendation. The model characterizes the interplay between user-item collaborative views and item multi-modal semantic views, aiming to construct a sophisticated model of long-term dependencies. In particular, we design a multi-scale dilated attention mechanism, which captures long-range dependencies through local and sparse block interactions from visual data. Additionally, SS-GANs based on assisted rotational loss is implemented to utilizes users’ historical interaction data to predict events that have not yet occurred, thus learning potential representations of users and items, which greatly improves the stability of feature embedding processing. Our extensive experiments on real-world datasets have shown that this method substantially improves performance over existing advanced baselines in multi-modal recommendation tasks, highlighting its potential and effectiveness.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-scale Dilated Attention Enhanced Multi-modal Self-supervised GANs for Recommendation

  • Yi Xue,
  • Qianqian Ren,
  • Hu Jin

摘要

Online multi-modal sharing platforms have necessitated the integration of diverse data types such as visual, textual, and auditory to form a holistic user preference profile in personalized recommendation systems. Current multi-modal recommendation methods extract features from user-item interactions but often depend on custom tags and face high computational costs, particularly in visual feature processing, and struggle with effectively modeling long-range user-item dependencies. To overcome these challenges, we introduce a novel approach, Multi-perspective Attention Enhanced Multi-modal Self-supervised GANs for Recommendation. The model characterizes the interplay between user-item collaborative views and item multi-modal semantic views, aiming to construct a sophisticated model of long-term dependencies. In particular, we design a multi-scale dilated attention mechanism, which captures long-range dependencies through local and sparse block interactions from visual data. Additionally, SS-GANs based on assisted rotational loss is implemented to utilizes users’ historical interaction data to predict events that have not yet occurred, thus learning potential representations of users and items, which greatly improves the stability of feature embedding processing. Our extensive experiments on real-world datasets have shown that this method substantially improves performance over existing advanced baselines in multi-modal recommendation tasks, highlighting its potential and effectiveness.