A Chinese natural speech complex emotion dataset based on emotion vector annotation method
摘要
Although Chinese speech emotion recognition has received increasing attention, existing datasets still have defects such as insufficient naturalness, unreliable annotation, and single pronunciation style, which seriously hinder research progress. To address these issues, this paper proposes a Chinese natural speech complex emotion dataset (CNSCED) to provide natural data resources for Chinese speech affective computing. CNSCED was curated from publicly available Chinese news and interview television programs, capturing authentic emotional expressions encountered in daily life. The dataset comprises 14 h of speech from 454 speakers of diverse ages, totaling 15,777 samples. Acknowledging the inherent complexity and ambiguity of natural emotions, we propose an emotion vector annotation method. This method utilizes a vector composed of six meta-emotion dimensions (anger, sadness, aroused, happiness, surprise, and fear) of different intensities to describe any single or complex emotional state. CNSCED released two subtasks: complex emotion classification and complex emotion intensity detection. In the experiment, we evaluated the CNSCED using deep neural network models and provided a baseline result. To the best of our knowledge, CNSCED is the first publicly available Chinese natural speech complex emotion dataset, which can be freely downloaded from https://github.com/wuxlxju/CNSCED.