Robust Automatic Speech Recognition Using Wavelet-Based Adaptive Wavelet Thresholding: A Review
摘要
Automatic speech recognition (ASR) is one of the most fascinating fields of research and the performance of ASR systems is most promising in a closed environment having negligible or zero background noise. However, the performance is not satisfactory if the spoken environment has a wide variety of noises. The performance of ASR systems is also influenced by the size of the vocabulary, speech sound unit, spoken environment, native language influences, transmission channel, emotional and health state of the speaker, age, speech corpus design, and its preprocessing and other challenges. The presence of noise in spoken speech sound significantly affects the ASR system performance out of all other challenges. So, in this paper, introduction to ASR and its history, human speech production and perception, ASR terminologies, the framework of ASR, and challenges associated with the design of ASR are discussed in detail with a focus on speech corpus design and its preprocessing. The traditional methods of speech enhancement based on the time domain and frequency domain are not able to handle the nonstationary noise that is present in the speech signals because of the fixed window duration; the wavelet transforms-based soft thresholding-based speech enhancement techniques, on the other hand, have the ability to handle nonstationary noise by applying windows of variable durations serve as an alternative powerful tool for preprocessing of the speech signals contaminated by additive Gaussian noise of various SNR levels. The speech signals are first decomposed into high- and low-frequency subbands and it is a well-known fact that the majority of the noise is present in the high-frequency bands. The wavelet-based Bayes shrink algorithm along with a combination of time-adaptive thresholds and soft thresholding technique is applied to the high-frequency subbands to remove the noise, and the performance of the soft thresholding method is compared with that of the state of art hard thresholding method. The experiments are carried out on the Kannada speech corpus consisting of continuous sentences, and the results are compared with those obtained on the benchmarking TIMIT dataset. The results convey that the proposed method performs better for SNR of various levels.