Building an Egyptian-Arabic Speech Corpus for Emotion Analysis Using Deep Learning
摘要
Emotionally intelligent Virtual Assistants (VAs) are increasingly gaining popularity, especially with the digitization of different life aspects. The focus of our work is to build VAs that can understand the emotional state of users from their Egyptian-Arabic speech. This requires the availability of large emotional datasets to be able to train accurate models. Available corpora include different languages and dialects. However, the Egyptian-Arabic dialect, in particular, shows a significant gap. The main contribution of this paper is to fill this gap by gathering a semi-natural Egyptian-Arabic dataset. The dataset includes six emotions: happiness, sadness, anger, neutral, surprise, and fear. To the best of the authors’ knowledge, it is considered as the first Egyptian-Arabic dataset to include surprise and fear emotions. Also, a Deep Learning (DL) model is introduced that is able to detect the first 4 emotions with average accuracies of 70.3% and 73% for an imbalanced dataset and a balanced dataset, respectively, and the first 5 emotions with average accuracies of 65% and 66% for an imbalanced dataset and a balanced dataset, respectively.