Enhanced artificial neural network-based SER model in low-resource Indian language
摘要
Speech emotion recognition (SER) is an emerging application that helps computers understand human intentions and preferences. However, the task of SER in low-resource Indian languages, such as Bengali, remains challenging due to content variations, dialectical shifts, acoustic variability, and speaker age variations. This study presents an artificial neural network (ANN)-based model for speech emotion recognition in the Bengali language. The model uses spectral features like short-time Fourier transform (STFT), chroma-VQT, Melspectrogram, and mel-frequency cepstral coefficient (MFCC) to understand and recognize the emotions people talk about. Four frequency analysis tools are used to convert time-domain signals into spectral features extracted from the BanglaSER dataset. It is reported that the proposed ANN model with MFCC spectral feature, compiled with Nadam optimizer, yields higher train and test accuracies compared to other spectral feature combinations. The proposed model is the first attempt to employ an ANN-based deep learning model in the Bengali SER task, generating significantly higher accuracies than conventional machine-learning-based models.