Gotta embed them all! - knowledge-aware recommendations fusing heterogeneous multimodal item embeddings
摘要
In this paper, we present a methodology to provide users with knowledge-aware recommendations based on the fusion of multimodal item embeddings. Our approach relies on the intuition that each modality (i.e., graph, text, video, images, etc.) emphasizes different characteristics and nuances of the items, so it is necessary that a comprehensive knowledge-aware recommender system (KARS) encodes and exploits all the different data sources that are available in a specific domain. Accordingly, we design a multimodal KARS architecture based on a deep neural network that: (a) learns a representation of each uni-modal feature (i.e., description, trailers, covers, audio signals, and so on) through an appropriate encoder; (b) exploits self-attention and cross-attention to fuse the different sources and refine the embeddings; (c) returns a prediction score which represents user’s interest in the item, which is finally used to generate a top-k recommendation list. In the evaluation, we carried out experiments against two datasets, and the results showed that our approach overcame several baselines for multimodal and knowledge-aware recommendations, thus confirming the intuitions behind this work.