错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Standardized fusion of phase features with MFCCs for speech recognition using LSTM

  • Usha Pagadala,
  • A Krishna Chaitanya

摘要

In this paper, we propose two phase feature extraction methods to improve the automatic speech recognition (ASR) performance. Existing approaches primarily focus on suppressing spurious spikes in the group delay (GD) function but often neglect the issue of large dynamic range variations. To address this limitation, we propose selective frequency inclusion group delay (SFIGD), which preserves essential vocal tract information by optimally selecting significant frequency components, and dynamic range alteration group delay (DRAGD), which reduces dynamic range to minimize information loss. Phase features obtained from these methods are evaluated individually and in fusion with Mel-frequency cepstral coefficients (MFCCs) using a long short-term memory (LSTM) network. Experiments are conducted on isolated spoken word recognition tasks using the AudioMNIST dataset for English digits and a custom Telugu dataset. Results demonstrate significant improvements in recognition accuracy, achieving up to \(96.54\%\) 96.54 % for English and \(88.05\%\) 88.05 % for Telugu in fused settings. The proposed methods outperform conventional phase features and MFCCs, highlighting their effectiveness for both high and low-resource languages.