Computational music emotion recognition (MER) is an important task that aims to recognize emotional content in music tracks. Understanding the emotional content of music can help tailor therapeutic interventions to specific emotional needs, potentially improving their effectiveness. Piano music is widely used in music therapy for depression, anxiety, and other mental health conditions. However, comparatively little work has been done on MER in piano music, which is a challenging task due to the unique difficulties in recognizing emotions in single-timbre music without vocals, especially on valence. In this paper, we propose a novel Double-Mix-Net model that utilizes the multi-layer feature mixing and multi-modal feature mixing methods to improve classification performance. Our model uses both the title of the YouTube music video and the music context to predict the emotion of a given music track. Additionally, we make full use of different granularity features through the multi-layer feature mixing method. Our proposed model achieved state-of-the-art accuracy performance of 73.8% on the Emopia dataset and is the first multi-modal network designed for piano emotion recognition. The results of our study highlight the importance of considering the music title and utilizing multi-layer feature mixing methods in piano music emotion recognition. This work helps the emotional understanding of piano music. Our proposed model can be applied in various domains such as music recommendation systems, film scoring, and game design to enhance the emotional impact of music.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Double-Mix-Net: A Multimodal Music Emotion Recognition Network with Multi-layer Feature Mixing

  • Peilin Li,
  • Kairan Chen,
  • Weixin Wei,
  • Jiahao Zhao,
  • Wei Li

摘要

Computational music emotion recognition (MER) is an important task that aims to recognize emotional content in music tracks. Understanding the emotional content of music can help tailor therapeutic interventions to specific emotional needs, potentially improving their effectiveness. Piano music is widely used in music therapy for depression, anxiety, and other mental health conditions. However, comparatively little work has been done on MER in piano music, which is a challenging task due to the unique difficulties in recognizing emotions in single-timbre music without vocals, especially on valence. In this paper, we propose a novel Double-Mix-Net model that utilizes the multi-layer feature mixing and multi-modal feature mixing methods to improve classification performance. Our model uses both the title of the YouTube music video and the music context to predict the emotion of a given music track. Additionally, we make full use of different granularity features through the multi-layer feature mixing method. Our proposed model achieved state-of-the-art accuracy performance of 73.8% on the Emopia dataset and is the first multi-modal network designed for piano emotion recognition. The results of our study highlight the importance of considering the music title and utilizing multi-layer feature mixing methods in piano music emotion recognition. This work helps the emotional understanding of piano music. Our proposed model can be applied in various domains such as music recommendation systems, film scoring, and game design to enhance the emotional impact of music.