<p>Understanding the interplay between rhythm and emotion is critical for advancing automatic music emotion recognition (MER). While previous works have made progress in modeling time–frequency features, most fail to explicitly capture frequency-domain rhythmic salience or exploit the joint structure of rhythm and emotion. In this paper, we propose a novel multi-task learning framework that integrates frequency-domain attention (ATFNet), local convolutional encoding, and bidirectional temporal modeling (BiLSTM) to perform rhythm type classification, rhythm intensity regression, and valence–arousal prediction simultaneously. ATFNet enhances rhythm-sensitive frequency bands by applying Fourier transforms and frequency-channel attention, while the convolutional module extracts local spectral dynamics, and BiLSTM captures long-range emotional dependencies. Furthermore, we introduce an adversarial recognition module to improve feature discrimination and generalization across music genres. We evaluate our model on the DEAM dataset under both the traditional DMER and personalized DMER (PDMER) settings. Results show that our method achieves the best performance across all metrics. In the DMER setting, our model attains a CCC of 0.368 and PCC of 0.633 for arousal prediction, and 0.098/0.114 for valence, significantly outperforming strong baselines such as DAMFF and CRNN. In the PDMER task, our model achieves CCC scores of 0.333 (arousal) and 0.078 (valence), with consistent gains in PCC and RMSE. These results demonstrate the robustness and effectiveness of our joint rhythm-emotion modeling framework for fine-grained music understanding.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Joint modeling of rhythm and emotion in music via frequency-domain attention and multi-task learning

  • Shengrui Liu,
  • Lei Ren,
  • Jiaxuan Zhao

摘要

Understanding the interplay between rhythm and emotion is critical for advancing automatic music emotion recognition (MER). While previous works have made progress in modeling time–frequency features, most fail to explicitly capture frequency-domain rhythmic salience or exploit the joint structure of rhythm and emotion. In this paper, we propose a novel multi-task learning framework that integrates frequency-domain attention (ATFNet), local convolutional encoding, and bidirectional temporal modeling (BiLSTM) to perform rhythm type classification, rhythm intensity regression, and valence–arousal prediction simultaneously. ATFNet enhances rhythm-sensitive frequency bands by applying Fourier transforms and frequency-channel attention, while the convolutional module extracts local spectral dynamics, and BiLSTM captures long-range emotional dependencies. Furthermore, we introduce an adversarial recognition module to improve feature discrimination and generalization across music genres. We evaluate our model on the DEAM dataset under both the traditional DMER and personalized DMER (PDMER) settings. Results show that our method achieves the best performance across all metrics. In the DMER setting, our model attains a CCC of 0.368 and PCC of 0.633 for arousal prediction, and 0.098/0.114 for valence, significantly outperforming strong baselines such as DAMFF and CRNN. In the PDMER task, our model achieves CCC scores of 0.333 (arousal) and 0.078 (valence), with consistent gains in PCC and RMSE. These results demonstrate the robustness and effectiveness of our joint rhythm-emotion modeling framework for fine-grained music understanding.