Multimodal Fusion-Based Hybrid CRNN Model for Emotion Prediction in Music
摘要
Many people believe that music is an expression of emotions. It has long been known how it generates emotional evaluation in individuals from various communities, cultures, and demographic groups; as a result, categorising it on the basis of emotions is, in fact, an exciting fundamental study field. In this study, we utilised MediaEval, a subset of DEAM, MER500 dataset, MediaEval dataset, which now has a total of 1797 audio files of different scales. The convolutional recurrent neural network (CRNN) is fed the spectrogram in order to gather the time-domain, frequency-domain, and sequence features of audio. The BiLSTM network receives low-level audio data concurrently in order to further extract the sequence information of audio features. The two feature sections are concatenated and fed into the Support Vector Machines classification function utilising the centre loss function in order to distinguish four musical emotions. Testing results show that the suggested approach outperforms existing techniques in terms of recognition accuracy (97.72%). The recommended method presents a novel, useful suggestion for the advancement of musical emotion identification.