A Multi-modal Short-Video Recommendation Method for Enhancing User Behaviors Based on Graph Contrastive Learning
摘要
Although existing short video recommendation methods have explored user-video interaction behaviors, they still suffer from data sparsity and weak supervision signals due to the scarcity of single or partial multi-behavior data (e.g., views, likes, and favorites). Furthermore, heterogeneous noise across behavior types, particularly the significantly higher noise ratio in “view” interactions, severely degrades the accuracy of multi-behavior short video recommendation systems. Therefore, we propose MBGCL, a multi-modal short-video recommendation method for enhancing user behaviors based on graph contrastive learning, to address these limitations. Firstly, the various modal data (text, image, and video) of short videos are dynamically fused to obtain more accurate short video representations. Secondly, the self-supervised contrastive learning method is introduced to learn the features of nodes from the perspective of behaviors, solving the problems of data scarcity and weak supervision signals. Finally, to reduce the problem caused by behavior noise, the behavior subgraph enhanced contrastive learning is introduced to counteract the negative transfer effect of behavior noise on the recommendation results. Through a large number of experiments on three public datasets (TikTok, KuaiRand, and MicroLens-100K), it is proved that our proposed model significantly outperforms the baseline models.