<p>To enhance the accuracy and expressiveness of emotional recognition in instrumental music, this study proposes a multi-layer music emotion recognition model. The model integrates Convolutional Neural Network (CNN), Bidirectional Gated Recurrent Unit (BiGRU), and Attention Mechanism, aiming to accurately capture complex emotional information in audio. First, the CNN is used to extract local emotional features of the audio. Then, the BiGRU is employed to model the contextual information of time series, strengthening the temporal continuity of emotional expression. Finally, the attention mechanism is introduced to dynamically focus on key emotional segments. To achieve multi-scale feature fusion, the model combines low-level audio features and high-level semantic features through weighted summation during the feature extraction stage. The experimental section is validated using three music emotion datasets, including two publicly available datasets Instrument Recognition in Musical Audio Signals (IRMAS) and Multitrack Dataset for Musical Audio (MedleyDB), as well as a large-scale dataset Database for Emotional Analysis of Music (DEAM), to comprehensively evaluate the performance, generalization ability, and robustness of the model. These datasets cover a large number of multi-category instrumental audio samples. The model is evaluated on three continuous emotional dimensions: Valence, Arousal, and Dominance. The experimental results show that the proposed model achieves Pearson correlation coefficients of 0.871, 0.832, and 0.784, respectively, which are better than those of the comparative models. In terms of Mean Squared Error (MSE), the values are 0.0187, 0.0208, and 0.0243, respectively, indicating higher prediction accuracy. In conclusion, the proposed fusion deep neural network model significantly improves the accuracy and generalization ability of emotional recognition in instrumental music. This study provides an effective method and practical inspiration for emotional modeling in complex musical environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Application of artificial intelligence CNN model in emotional recognition of instrumental music

  • Yang Liu

摘要

To enhance the accuracy and expressiveness of emotional recognition in instrumental music, this study proposes a multi-layer music emotion recognition model. The model integrates Convolutional Neural Network (CNN), Bidirectional Gated Recurrent Unit (BiGRU), and Attention Mechanism, aiming to accurately capture complex emotional information in audio. First, the CNN is used to extract local emotional features of the audio. Then, the BiGRU is employed to model the contextual information of time series, strengthening the temporal continuity of emotional expression. Finally, the attention mechanism is introduced to dynamically focus on key emotional segments. To achieve multi-scale feature fusion, the model combines low-level audio features and high-level semantic features through weighted summation during the feature extraction stage. The experimental section is validated using three music emotion datasets, including two publicly available datasets Instrument Recognition in Musical Audio Signals (IRMAS) and Multitrack Dataset for Musical Audio (MedleyDB), as well as a large-scale dataset Database for Emotional Analysis of Music (DEAM), to comprehensively evaluate the performance, generalization ability, and robustness of the model. These datasets cover a large number of multi-category instrumental audio samples. The model is evaluated on three continuous emotional dimensions: Valence, Arousal, and Dominance. The experimental results show that the proposed model achieves Pearson correlation coefficients of 0.871, 0.832, and 0.784, respectively, which are better than those of the comparative models. In terms of Mean Squared Error (MSE), the values are 0.0187, 0.0208, and 0.0243, respectively, indicating higher prediction accuracy. In conclusion, the proposed fusion deep neural network model significantly improves the accuracy and generalization ability of emotional recognition in instrumental music. This study provides an effective method and practical inspiration for emotional modeling in complex musical environments.