Standardized fusion of phase features with MFCCs for speech recognition using LSTM
摘要
In this paper, we propose two phase feature extraction methods to improve the automatic speech recognition (ASR) performance. Existing approaches primarily focus on suppressing spurious spikes in the group delay (GD) function but often neglect the issue of large dynamic range variations. To address this limitation, we propose selective frequency inclusion group delay (SFIGD), which preserves essential vocal tract information by optimally selecting significant frequency components, and dynamic range alteration group delay (DRAGD), which reduces dynamic range to minimize information loss. Phase features obtained from these methods are evaluated individually and in fusion with Mel-frequency cepstral coefficients (MFCCs) using a long short-term memory (LSTM) network. Experiments are conducted on isolated spoken word recognition tasks using the AudioMNIST dataset for English digits and a custom Telugu dataset. Results demonstrate significant improvements in recognition accuracy, achieving up to