错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Toward an emotion efficient architecture based on the sound spectrum from the voice of Portuguese speakers

  • Geraldo P. Rocha Filho,
  • Rodolfo I. Meneguette,
  • Fábio Lúcio Lopes de Mendonça,
  • Liriam Enamoto,
  • Gustavo Pessin,
  • Vinícius P. Gonçalves

摘要

One of the main challenges in the process of recognizing emotion through the voice are related to the specific characteristics of an individual’s sound spectrum, such as accent and speech rhythm, as well as regionalism and wide variability of spoken phrases. Despite efforts to propose emotion recognition models, providing an increase in accuracy in classifying emotion in a specialized way is an open research question. Faced with these challenges, this work proposes DEEP (DEtection of voice Emotion in Portuguese language), an architecture for detecting voice emotion based on patterns present in the sound spectrum generated by the voice of Brazilian Portuguese speakers. DEEP recognizes each emotion by using a set of specialist Convolutional Neural Networks that receive as input the features extracted from the sound spectrum. With this, DEEP aims to specialize each emotion to increase the rate of correct answers and adapt to different tones and voice conditions that may occur in everyday life. Our results show that DEEP outperforms the emotion recognition measures of other state of art techniques for all evaluated scenarios.