错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Application of Multimodal Machine Learning for Image Recommendation Systems

  • Mikhail Foniakov,
  • Anatoly Bardukov,
  • Ilya Makarov

摘要

In the era of information overload, making decisions about various aspects of our lives has become increasingly challenging. The rise of the internet and online platforms has provided access to vast amounts of information, including recommendations based on user preferences. This paper focuses on the development of a unique multimodal recommender system for images, leveraging machine learning and deep learning techniques. The results demonstrate the effectiveness of a complex recommender system, with an average accuracy of 0.814. The system outperforms algorithms based solely on CLIP embeddings or BERT vectors, showcasing the advantages of incorporating multiple modalities. The strengths of the model lie in its robust recognition of CLIP vectors and efficient processing of image features. However, there is room for improvement in text features and embeddings, suggesting the need for more detailed textual information and additional data sources. The innovation of this system is to utilize both image data and text descriptions to provide personalized recommendations. Images are vectorized and combined with relevant metrics, while text descriptions are obtained through object recognition or dataset text. The model incorporates various data, such as location, creation date, and event information, to enhance the recommendation process. Overall, this study highlights the potential of multimodal recommender systems for enhancing recommendation quality and providing users with a personalized experience. Ongoing efforts to refine the model with diverse data and parameter optimization can further improve its performance.