Enhancement and reconstruction of dysphonic Kannada speech using LSTM and convolution network
摘要
Dysphonia is a prevalent speech disorder that manifests as irregularities in vocal quality, which presents significant challenges to effective communication. This disorder not only hampers day-to-day interactions but also profoundly affects an individual’s quality of life. Despite its impact, there is a limited focus on addressing dysphonia for non-English languages. This work introduces a methodology for the enhancement of dysphonic speech in Kannada, one of the most widely spoken languages in India. The proposed solution leverages advanced techniques in signal processing and machine learning to restore clarity and intelligibility to dysphonic speech, thereby facilitating better communication for affected individuals. The dataset consists of regularly speaking Kannada sentences recorded from dysphonic subjects in low noise environment. Additive noise present while recording is reduced using spectral subtraction. Noise reduced dysphonic speech is enhanced and reconstructed using deep learning methods like long short-term memory (LSTM) and convolution neural network (CNN). The outcome of the methods is analyzed using quantitative evaluation.