Multilingual Emotion Recognition from Continuous Speech Using Transfer Learning
摘要
The ability to recognize emotions from audio is currently in high demand across various fields, such as Intelligence Services, Journalism, and Security. The COVID-19 pandemic has resulted in a shift to online meetings, conferences, and interrogations, making automated emotion detection from audio crucial. It is important to recognize emotions in languages other than English as well, due to the multilingual population. This paper presents a real-time emotion detection system that utilizes both acoustic and linguistic features of audio to predict emotions. Real time Microcontroller is used to capture videos to be used by the proposed Hubert based model. To train the deep learning model, audio files are extracted from videos in the PU Dataset and the RAVDESS Dataset. The transfer learning approach is used to retrain the proposed model for specific languages separately. For robust training, multi lingual PU Dataset is used having videos of individuals from various ethnicities with different linguistic features and accents. The best test accuracy achieved for the English, Hindi, and Punjabi languages are 95.95%, 88.6364%, and 89.70%, respectively.