The integration of emotion detection with music recommendation systems has the potential to enhance user experience by delivering music tailored to an individual’s emotional state. This paper provides a comparative study of two machine learning models used in a smart music player system. The first model uses Haar Cascade for face detection and a Convolutional Neural Network (CNN) for emotion classification, whereas the second model uses state-of-the-art methods like RetinaFace for face detection and a combination of ResNet and EfficientNet for emotion recognition. The dataset used for the comparison is FER2013. Both models connect the recognized emotions to a pre-curated music database for personalized music suggestions. Performance is evaluated based on accuracy, detection time, and responsiveness to changing lighting conditions and angles of the face. Initial results show that the newer model has better accuracy and detection speed, and thus is more applicable to real-time implementations. Although the limitation of availability of diverse dataset and hardware such as GPU’s for better performance is always there. This paper brings out improvements in emotion recognition methods and their importance for user-oriented multimedia applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative Analysis of Machine Learning Models for Emotion-Based Music Recommendation Systems

  • Shreyas Gahtori,
  • Kshitij Chitransh,
  • Deependra Singh,
  • Sagar Uniyal

摘要

The integration of emotion detection with music recommendation systems has the potential to enhance user experience by delivering music tailored to an individual’s emotional state. This paper provides a comparative study of two machine learning models used in a smart music player system. The first model uses Haar Cascade for face detection and a Convolutional Neural Network (CNN) for emotion classification, whereas the second model uses state-of-the-art methods like RetinaFace for face detection and a combination of ResNet and EfficientNet for emotion recognition. The dataset used for the comparison is FER2013. Both models connect the recognized emotions to a pre-curated music database for personalized music suggestions. Performance is evaluated based on accuracy, detection time, and responsiveness to changing lighting conditions and angles of the face. Initial results show that the newer model has better accuracy and detection speed, and thus is more applicable to real-time implementations. Although the limitation of availability of diverse dataset and hardware such as GPU’s for better performance is always there. This paper brings out improvements in emotion recognition methods and their importance for user-oriented multimedia applications.