A Comprehensive Performance Evaluation of Whisper Models in Dysarthric Speech Recognition
摘要
In this paper, we explore the application of the state-of-the-art Whisper model for automatic speech recognition (ASR) of dysarthric speech. Dysarthric speech presents significant challenges for conventional ASR systems due to factors such as severely affected speech intelligibility, speaker variations, and data scarcity. The Whisper model, trained on extensive multilingual datasets, has shown remarkable results for typical speech, but its performance in dysarthric speech remains underexplored. This study evaluates the Whisper model’s different versions (English-only and multilingual) and various model sizes (tiny, base, small, and medium) in typical speech settings. Our empirical results demonstrate that the medium-sized multilingual Whisper model achieves outstanding performance on the TORGO dataset, with a CER of 6.83% and WRA of 88.86% for isolated speech and a CER of 1.34% and WER of 3.16% for continuous speech. Furthermore, the medium-sized English-only model outperforms its multilingual counterparts on the UASpeech dataset across all intelligibility categories, achieving an average CER of 36.14% and a WRA of 59.96%.