<p>Pathological speech recognition presents the unique challenge of understanding and transcribing speech patterns that deviate from the normal due to medical conditions or impairments. Effective clinical tools, diagnostic procedures, and assistive technologies for individuals with communication difficulties depend on the development of robust models for this task. The scarcity of datasets containing speech from individuals with various pathologies poses a significant hurdle in training comprehensive models. This work investigates the impact of signal representations and data augmentation techniques in building accurate pathological speech recognition models. Initially, short-time Fourier transform (STFT) was used to convert speech into spectrograms for model training and evaluation. While performing adequately on normal speech, accuracy declined significantly with pathological speech data, highlighting the limitations of a small pathological dataset. To address this, multiple data augmentation strategies were employed. First, several variations of the STFT spectrogram were generated, including Mel-scaled STFT (STFT-Mel), STFT mel Perceptual Enhancement (STFT-Mel-PE), Constant-Q Transform (CQT), and single frequency filtered (SFF) spectrogram. Then, further augmented the dataset by manipulating spectrogram parameters like frame size and bin size. These techniques led to substantial expansion in the signal representation and noticeable performance improvements with Kannada word recognition tasks for hearing- impaired speech. A drastic improvement in the classifier performance from 23 to 82% was observed with the proposed data augmentation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Isolated Word Classification of Hearing Impaired Speech Using Time–Frequency Representations

  • Y. A. Goutham,
  • T. S. Himasagar,
  • Veena Karjigi,
  • H. M. Chandrashekar,
  • N. Sreedevi

摘要

Pathological speech recognition presents the unique challenge of understanding and transcribing speech patterns that deviate from the normal due to medical conditions or impairments. Effective clinical tools, diagnostic procedures, and assistive technologies for individuals with communication difficulties depend on the development of robust models for this task. The scarcity of datasets containing speech from individuals with various pathologies poses a significant hurdle in training comprehensive models. This work investigates the impact of signal representations and data augmentation techniques in building accurate pathological speech recognition models. Initially, short-time Fourier transform (STFT) was used to convert speech into spectrograms for model training and evaluation. While performing adequately on normal speech, accuracy declined significantly with pathological speech data, highlighting the limitations of a small pathological dataset. To address this, multiple data augmentation strategies were employed. First, several variations of the STFT spectrogram were generated, including Mel-scaled STFT (STFT-Mel), STFT mel Perceptual Enhancement (STFT-Mel-PE), Constant-Q Transform (CQT), and single frequency filtered (SFF) spectrogram. Then, further augmented the dataset by manipulating spectrogram parameters like frame size and bin size. These techniques led to substantial expansion in the signal representation and noticeable performance improvements with Kannada word recognition tasks for hearing- impaired speech. A drastic improvement in the classifier performance from 23 to 82% was observed with the proposed data augmentation.