Noise-Robust Punjabi Speech Enhancement Using LSTM and Feature Fusion Techniques
摘要
This paper presents a novel and robust deep learning-based Speech Enhancement Framework for Punjabi (SEFP) that integrates dual Long Short-Term Memory (LSTM) networks with a Kalman Filter for effective noise suppression. The framework is designed to enhance the intelligibility of Punjabi speech in noisy environments, a significant challenge for low-resource tonal languages. The first LSTM model estimates the clean magnitude spectrum from noisy acoustic features, while the second predicts Line Spectrum Frequencies (LSFs), which are subsequently transformed into Linear Predictive Coding (LPC) coefficients for time-domain enhancement using a Kalman Filter. The system processes multiple acoustic and prosodic features—namely MFCC, GFCC, BFCC, and pitch—to exploit complementary information through feature fusion. To ensure the scientific validity of results, the model was evaluated on two datasets: a newly developed Punjabi speech corpus and the publicly available Common Voice Punjabi dataset. Additionally, Word Error Rate (WER) was computed using a standardized Kaldi-based ASR system with consistent decoding parameters. Comparative evaluations against classical baselines (Wiener filtering, DNN-IRM) and fivefold cross-validation were conducted to assess generalizability. The optimal feature combination (Pitch + MFCC + GFCC + BFCC) achieved the lowest average WER of 20.86% across four noise types from the NOISEX-92 dataset—validating the system’s effectiveness and reliability for practical deployment in noisy speech scenarios.