Acquiring Knowledge for Mimicking Dysarthric Speech by Incorporating Its Features into Synthetic Speech
摘要
The purpose of this study is twofold. First of all, it is to conduct in-depth analyses, allowing the extraction of features associated with dysfunctional speech, and in particular, dysarthria. Temporal and spectral speech signal features are investigated. This analytical approach results in a set of features that corresponds best to dysarthria. The results are shown in the form of stationary analyses regarding a sequence of dysarthric utterances to highlight detected changes. For the purpose of parameter comparison, Kernel Density Estimation (KDE) was exploited. In addition, statistical analysis of some specific speech parameters is performed. The second purpose of this study is to propose techniques that may be employed to synthesize normal speech patterns with the most relevant dysarthria features to create dysfunctional speech. Their outcome is examined both by objective measures, such as WER (Word Error Rate) and CER (Character Error Rate), employing transformer-based speech recognition. The summary of the experiments includes conclusions and plans for future research studies related to automatic recognition of dysarthric speech.