Enhancing Recommendation Systems: A Comparative Analysis of Multimodal Feature Representations
摘要
Recent studies have shown that multimodal feature representations improve recommendation accuracy and help address the cold-start problem. These representations are achieved through deep learning-based fusion of various modalities or knowledge graph-based text encoding. However, training multimodal recommendation models still relies heavily on labelled data, which is labour-intensive and costly to obtain. In this paper, we conduct a comparative analysis of multidimensional recommender systems, exploring different combinations of modalities. Researchers are increasingly focusing on multimodal information to enhance recommendation systems. By leveraging rich contextual data such as textual descriptions, metadata, audio, and visual content, multimodal recommendation models can overcome the limitations of traditional collaborative filtering techniques.