Feature centric based deep learning approach for music mood recognition with HuBERT transformer model
摘要
In this study, music mood classification is explored using advanced deep learning and transformer-based models to accurately predict the emotional content of music. Music plays a crucial role in human life, influencing emotions, behaviors, and mental states. Accurately classifying the mood of music is essential for applications such as music recommendation systems, emotional intelligence in AI, and mental health monitoring. With the growing impact of Artificial Intelligence (AI) in emotion recognition, sentiment analysis has become a vital research area, contributing to fields like social media monitoring, customer experience analysis, and personalized content delivery. We employ a state-of-the-art transformer-based model (HuBERT) in comparison with deep learning models (ConvFormer, LSTM) and pre-trained models (YAMNet) to evaluate their effectiveness in music mood classification. The study is conducted on a publicly available dataset containing five mood labels such as Aggressive, Happy, Dramatic, Sad, and Romantic with 500 audio files per category. To analyze and extract meaningful patterns from the dataset, advanced different five audio features such as Short-Time Fourier Transform (STFT), and Mel-Frequency Cepstral Coefficients (MFCC) are utilized. To ensure robust model evaluation, an 80–20 holdout split is applied. Our experimental results indicate that HuBERT achieves the highest classification accuracy of 95%, outperforming both traditional deep learning models and pre-trained architecture. This study provides a comprehensive evaluation of machine learning approaches for music mood classification, demonstrating the potential of transformer-based models in enhancing AI-driven sentiment analysis. The findings contribute to the development of intelligent music recommendation systems, emotional AI applications, and advancements in affective computing.