错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Analysis of the Impact Using Pre-emphasis Filter, Unvoiced Sounds, Frame Size and Feature Vector Size on Human Emotion Recognition by Voice and Machine Learning

  • Alain Manzo-Martínez,
  • Raydesel Sánchez,
  • Fernando Gaxiola,
  • Fernando Martínez-Reyes,
  • Raymundo Cornejo-García

摘要

The area of human emotion recognition by voice is a very active research area in recent years. This area has applications in human - computer interaction, robotics, mobile services, call centers, computer games, psychological evaluation, among others. In this work, we are interested in systems that involve speech signal as a media for recognizing emotions. These systems are complex due to the structural analysis that they carry out to the signal in time and frequency domains. Therefore, it is important to make an analysis to the signal about different factors, such as, frame size, feature vector size, the use, or not, of pre-emphasis filter and eliminating or not, the unvoiced sounds. The experiments were carried out considering two databases, EMODB and EMOVO which belong to the German and Italian languages, respectively. In addition, we used three audio features: (a) Mel Frequency Cepstrals Coefficients (MFCC), (b) Relative Spectra-MFCC, and Multiband Spectral Entropy Signature. For the classification stage, we used Multilayer Perceptrons (MLP), Support Vector Machines and k-Nearest Neighbors. The results let us conclude that the unvoiced sounds and the pre-emphasis filter must be ruled out to characterize the signal and to get better classification rates, if the German language is used. In respect to the Italian language, the results indicate that both elements are necessary since both provide robustness to the feature extraction process by using this language. The rates reached for the EMODB and EMOVO databases with the classification algorithms were 82.0% and 77.0%, respectively.