Speech emotion recognition using graph convolutional networks
摘要
Speech Emotion Recognition (SER) is a crucial area within the field of affective computing. This study proposes the use of Graph Convolutional Networks (GCNs) in SER by exploring temporal proximity. First, we introduce the GCN architecture for sequential speech data. Using a temporal graph, we initialize the edge connections on the basis of temporal proximity. Second, we propose to improve the representation of temporal dependencies by adding edge optimization using a two-stage Genetic Algorithm (GA). The dimension of the parameter matrix is split and merged into a secondary search in the neighboring area. Finally, experimental results are achieved based on three public available databases, EMO-DB, SAVEE and RAVDESS. Our proposed model not only improves the accuracy of emotion recognition, but also offers new understanding of temporal emotional dependencies in speech communication.