The purpose of this study is twofold. First of all, it is to conduct in-depth analyses, allowing the extraction of features associated with dysfunctional speech, and in particular, dysarthria. Temporal and spectral speech signal features are investigated. This analytical approach results in a set of features that corresponds best to dysarthria. The results are shown in the form of stationary analyses regarding a sequence of dysarthric utterances to highlight detected changes. For the purpose of parameter comparison, Kernel Density Estimation (KDE) was exploited. In addition, statistical analysis of some specific speech parameters is performed. The second purpose of this study is to propose techniques that may be employed to synthesize normal speech patterns with the most relevant dysarthria features to create dysfunctional speech. Their outcome is examined both by objective measures, such as WER (Word Error Rate) and CER (Character Error Rate), employing transformer-based speech recognition. The summary of the experiments includes conclusions and plans for future research studies related to automatic recognition of dysarthric speech.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Acquiring Knowledge for Mimicking Dysarthric Speech by Incorporating Its Features into Synthetic Speech

  • Tomasz Piernicki,
  • Gražina Korvel,
  • Bożena Kostek

摘要

The purpose of this study is twofold. First of all, it is to conduct in-depth analyses, allowing the extraction of features associated with dysfunctional speech, and in particular, dysarthria. Temporal and spectral speech signal features are investigated. This analytical approach results in a set of features that corresponds best to dysarthria. The results are shown in the form of stationary analyses regarding a sequence of dysarthric utterances to highlight detected changes. For the purpose of parameter comparison, Kernel Density Estimation (KDE) was exploited. In addition, statistical analysis of some specific speech parameters is performed. The second purpose of this study is to propose techniques that may be employed to synthesize normal speech patterns with the most relevant dysarthria features to create dysfunctional speech. Their outcome is examined both by objective measures, such as WER (Word Error Rate) and CER (Character Error Rate), employing transformer-based speech recognition. The summary of the experiments includes conclusions and plans for future research studies related to automatic recognition of dysarthric speech.