Enhanced perceptual wavelet packet features for spontaneous Kannada sentence recognition under uncontrolled conditions
摘要
A Robust Automatic Speech Recognition (ASR) system is proposed through a hybrid combination of Perceptual Wavelet Packet features, Deep Neural Network-Hidden Markov Model (DNN-HMM) acoustic models, and n-gram language models for recognizing Spontaneously spoken Kannada sentences. Two Experiments are performed on own Kannada speech corpus and the widely recognized TIMIT corpus. First one is the experiment on the clean speech datasets and the second one is on the speech degraded by noise data chosen from Noise-US dataset. The proposed system performance is evaluated through the metric Word Error Rate (WER). A comparative study on the performance analysis is presented. The results of the experimental investigations achieved an average reduction in Word Error Rate (WER) of 1.8% over Mel Frequency Cepstral Coefficients (MFCCs) and 2.1% over Perceptual Linear Prediction (PLP) features.