<p>Analysis of pathological speech is useful in many applications like intelligibility assessment and processing the same to improve the intelligibility. Phase is rarely used to analyse pathological speech. This work investigates the use of modified group delay, a phase based representation to assess the intelligibility of dysarthric speech from Universal Access database. Dysarthric speech is represented using short-time Fourier transform and Constant-Q transform. Spectrograms of real and imaginary components as well as magnitude and phase components contributed in different ways to the intelligibility assessment of dysarthric speech. Further, combination of magnitude and phase using product of power and modified group delay spectrogram is found to perform better than product of real and imaginary components. Constant-Q spectrogram derived from product of power and modified group delay when fed to convolutional neural network was able to assess the intelligibility of dysarthric speech with an accuracy of 95.5%.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Significance of magnitude and phase components for intelligibility assessment of pathological speech

  • H. M. Chandrashekar,
  • Veena Karjigi,
  • N. Sreedevi

摘要

Analysis of pathological speech is useful in many applications like intelligibility assessment and processing the same to improve the intelligibility. Phase is rarely used to analyse pathological speech. This work investigates the use of modified group delay, a phase based representation to assess the intelligibility of dysarthric speech from Universal Access database. Dysarthric speech is represented using short-time Fourier transform and Constant-Q transform. Spectrograms of real and imaginary components as well as magnitude and phase components contributed in different ways to the intelligibility assessment of dysarthric speech. Further, combination of magnitude and phase using product of power and modified group delay spectrogram is found to perform better than product of real and imaginary components. Constant-Q spectrogram derived from product of power and modified group delay when fed to convolutional neural network was able to assess the intelligibility of dysarthric speech with an accuracy of 95.5%.