In this paper, we explore the application of the state-of-the-art Whisper model for automatic speech recognition (ASR) of dysarthric speech. Dysarthric speech presents significant challenges for conventional ASR systems due to factors such as severely affected speech intelligibility, speaker variations, and data scarcity. The Whisper model, trained on extensive multilingual datasets, has shown remarkable results for typical speech, but its performance in dysarthric speech remains underexplored. This study evaluates the Whisper model’s different versions (English-only and multilingual) and various model sizes (tiny, base, small, and medium) in typical speech settings. Our empirical results demonstrate that the medium-sized multilingual Whisper model achieves outstanding performance on the TORGO dataset, with a CER of 6.83% and WRA of 88.86% for isolated speech and a CER of 1.34% and WER of 3.16% for continuous speech. Furthermore, the medium-sized English-only model outperforms its multilingual counterparts on the UASpeech dataset across all intelligibility categories, achieving an average CER of 36.14% and a WRA of 59.96%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Comprehensive Performance Evaluation of Whisper Models in Dysarthric Speech Recognition

  • Satwinder Singh,
  • Zihan Zhong,
  • Qianli Wang,
  • Clarion Mendes,
  • Mark Hasegawa-Johnson,
  • Waleed Abdulla,
  • Seyed Reza Shahamiri

摘要

In this paper, we explore the application of the state-of-the-art Whisper model for automatic speech recognition (ASR) of dysarthric speech. Dysarthric speech presents significant challenges for conventional ASR systems due to factors such as severely affected speech intelligibility, speaker variations, and data scarcity. The Whisper model, trained on extensive multilingual datasets, has shown remarkable results for typical speech, but its performance in dysarthric speech remains underexplored. This study evaluates the Whisper model’s different versions (English-only and multilingual) and various model sizes (tiny, base, small, and medium) in typical speech settings. Our empirical results demonstrate that the medium-sized multilingual Whisper model achieves outstanding performance on the TORGO dataset, with a CER of 6.83% and WRA of 88.86% for isolated speech and a CER of 1.34% and WER of 3.16% for continuous speech. Furthermore, the medium-sized English-only model outperforms its multilingual counterparts on the UASpeech dataset across all intelligibility categories, achieving an average CER of 36.14% and a WRA of 59.96%.