Enhancing Cross-Cultural Speech Emotion Recognition: A Comparative Study of Gujarati and English Datasets
摘要
Our study on Speech Emotion Recognition (SER) focuses on understanding emotions in Gujarati speech. We utilized the CREMA-D dataset for SER, as resources in Gujarati are limited. Cross-corpus analysis involved using both Gujarati and English databases due to similarities in cues. Testing Gujarati audios on an English SER model helped guide our decisions. Our study employed pre-existing models and refined them based on outcomes. Rigorous testing with various algorithms identified the most effective one for detecting emotions in Gujarati speech. The model learns to discern emotions from audio data, enhancing accuracy and robustness. We aim to integrate emotion recognition into daily life for personalized virtual assistants, mental health tools, and innovative applications.