<p>The objective of hyperkinetic dysarthria speech recognition is to translate spoken language into text, that has potential uses in fields such as healthcare, safety, and the automobile industry. In the education sector, it can assist students with speech difficulties in participating in classroom activities, completing assignments, and communicating with teachers and peers. Integrated techniques are becoming more popular in recent years, and they perform effectively wherever resources are limited. However, they are infrequently employed in East Slavic languages such as Russian language. In this research, a new hyperkinetic dysarthria dataset has been introduced, consisting of 6&#xa0;h, 26&#xa0;min, and 15&#xa0;s of recorded speech. Also, a novel feature extractor known as deep cascade convolution (DCC) architecture has been proposed to enhance the performance of disorder speech recognition. This extractor leverages convolution kernels of different sizes to gather and blend information from various scales. In addition, integrated technique for improving attention mechanism have been developed by using the connectionist temporal classification (CTC) objective function in training and the RNNTextGen language model in the decoding stage. The suggested technique enhances model convergence and improves hyperkinetic dysarthria speech recognition. Compared with the baseline model, the character error rate (CER) and word error rate (WER) on the test dataset decreased by 2.39% and 2.97% respectively. The findings indicate that proposed approach is comparable to the complex end-to-end systems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing hyperkinetic dysarthria speech recognition through deep cascade convolution and integrated attention mechanism

  • Antor Mahamudul Hashan

摘要

The objective of hyperkinetic dysarthria speech recognition is to translate spoken language into text, that has potential uses in fields such as healthcare, safety, and the automobile industry. In the education sector, it can assist students with speech difficulties in participating in classroom activities, completing assignments, and communicating with teachers and peers. Integrated techniques are becoming more popular in recent years, and they perform effectively wherever resources are limited. However, they are infrequently employed in East Slavic languages such as Russian language. In this research, a new hyperkinetic dysarthria dataset has been introduced, consisting of 6 h, 26 min, and 15 s of recorded speech. Also, a novel feature extractor known as deep cascade convolution (DCC) architecture has been proposed to enhance the performance of disorder speech recognition. This extractor leverages convolution kernels of different sizes to gather and blend information from various scales. In addition, integrated technique for improving attention mechanism have been developed by using the connectionist temporal classification (CTC) objective function in training and the RNNTextGen language model in the decoding stage. The suggested technique enhances model convergence and improves hyperkinetic dysarthria speech recognition. Compared with the baseline model, the character error rate (CER) and word error rate (WER) on the test dataset decreased by 2.39% and 2.97% respectively. The findings indicate that proposed approach is comparable to the complex end-to-end systems.