<p>In this paper, we present a methodology to provide users with <i>knowledge-aware recommendations</i> based on the fusion of multimodal item embeddings. Our approach relies on the intuition that each <i>modality</i> (<i>i.e.,</i> graph, text, video, images, etc.) emphasizes different characteristics and nuances of the items, so it is necessary that a comprehensive knowledge-aware recommender system (KARS) encodes and exploits all the different data sources that are available in a specific domain. Accordingly, we design a multimodal KARS architecture based on a deep neural network that: <i>(a)</i> learns a representation of each <i>uni-modal</i> feature (<i>i.e.,</i> description, trailers, covers, audio signals, and so on) through an appropriate encoder; <i>(b)</i> exploits <i>self-attention</i> and <i>cross-attention</i> to fuse the different sources and refine the embeddings; <i>(c)</i> returns a prediction score which represents user’s interest in the item, which is finally used to generate a<i> top-k</i> recommendation list. In the evaluation, we carried out experiments against two datasets, and the results showed that our approach overcame several baselines for multimodal and knowledge-aware recommendations, thus confirming the intuitions behind this work.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Gotta embed them all! - knowledge-aware recommendations fusing heterogeneous multimodal item embeddings

  • Giuseppe Spillo,
  • Elio Musacchio,
  • Cataldo Musto,
  • Marco de Gemmis,
  • Pasquale Lops,
  • Giovanni Semeraro

摘要

In this paper, we present a methodology to provide users with knowledge-aware recommendations based on the fusion of multimodal item embeddings. Our approach relies on the intuition that each modality (i.e., graph, text, video, images, etc.) emphasizes different characteristics and nuances of the items, so it is necessary that a comprehensive knowledge-aware recommender system (KARS) encodes and exploits all the different data sources that are available in a specific domain. Accordingly, we design a multimodal KARS architecture based on a deep neural network that: (a) learns a representation of each uni-modal feature (i.e., description, trailers, covers, audio signals, and so on) through an appropriate encoder; (b) exploits self-attention and cross-attention to fuse the different sources and refine the embeddings; (c) returns a prediction score which represents user’s interest in the item, which is finally used to generate a top-k recommendation list. In the evaluation, we carried out experiments against two datasets, and the results showed that our approach overcame several baselines for multimodal and knowledge-aware recommendations, thus confirming the intuitions behind this work.