Cross-view Contrastive Learning Enhanced Heterogeneous Graph Networks for Multi-modal Recipe Recommendation
摘要
Today’s recommender systems pay less attention to the synergistic effects of multi-modal information, such as recipe text, images, and relational data. Meanwhile, GNN-based recipe recommendation systems suffer from sparse supervised signals, significantly degrading their actual performance. In this paper, we propose a cross-view contrastive learning enhanced heterogeneous graph neural network (CCHGN) for multi-modal recipe recommendation. Specifically, we adopt a user-recipe-ingredient heterogeneous graph (URI-Graph) and integrate visual, textual, and relational information into the recipe embedding through a hierarchical attention GNN. Moreover, we extract the interaction between ingredients through an ingredient set transformer, and propose a multi-level cross-view contrastive learning mechanism for URI-Graph. We split the URI-Graph into local levels of user-recipe graph as collaborative view, recipe-recipe graph as similar view, and recipe-ingredient graph as attribute view. Extensive experiments show that CCHGN outperforms the state-of-the-art methods in recipe recommendation.