Emotional features of speech signals are one of the keys to human-computer interaction. However, there are still great difficulties and chances to extract emotional features. There is also great controversy regarding the part of signal preprocessing. This study divides the speech signal into small frames that overlap with a portion of the previous frame and adopts an improved empirical mode decomposition (EMD) based feature extraction method. The aim is to find the most suitable framing method. Each frame signal is processed by an improved EMD to generate a set of intrinsic mode functions (IMFs). Multidimensional features are extracted by calculating the central frequency and energy intensity of each IMF, and subsequently processing the center frequency of each IMF. Specifically, we focus on the top three IMFs in terms of energy intensity. Based on the improved algorithm, we investigate the effects of different frame lengths and frame shifts on the recognition rates of three emotion classifications: happy, angry, and sad. We find that the proposed method can reach the highest recognition rate when we use a 30 ms frame length with a 25% frame shift to separate the signals.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Frame Optimization in Speech Emotion Recognition Based on Improved EMD and SVM Algorithms

  • Chuan-Jie Guo,
  • Shu-Ya Jin,
  • Yu-Zhe Zhang,
  • Chi-Yuan Ma,
  • Muhammad Adeel,
  • Zhi-Yong Tao

摘要

Emotional features of speech signals are one of the keys to human-computer interaction. However, there are still great difficulties and chances to extract emotional features. There is also great controversy regarding the part of signal preprocessing. This study divides the speech signal into small frames that overlap with a portion of the previous frame and adopts an improved empirical mode decomposition (EMD) based feature extraction method. The aim is to find the most suitable framing method. Each frame signal is processed by an improved EMD to generate a set of intrinsic mode functions (IMFs). Multidimensional features are extracted by calculating the central frequency and energy intensity of each IMF, and subsequently processing the center frequency of each IMF. Specifically, we focus on the top three IMFs in terms of energy intensity. Based on the improved algorithm, we investigate the effects of different frame lengths and frame shifts on the recognition rates of three emotion classifications: happy, angry, and sad. We find that the proposed method can reach the highest recognition rate when we use a 30 ms frame length with a 25% frame shift to separate the signals.