错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dysphonia Diagnosis Using Self-supervised Speech Models in Mono and Cross-Lingual Settings

  • Dosti Aziz,
  • Dávid Sztahó

摘要

Voice disorders like dysphonia can significantly impact a person’s quality of life, so proper diagnostic methods are crucial. Previous approaches have primarily used datasets of a single language without considering language independence. This study investigates the effectiveness of self-supervised (SS) speech representation models in both mono- and cross-lingual settings to determine their ability to perform language-independent dysphonia detection. Four recent SS models, namely Wav2vec2.0, WavLM, HuBert, and Data2vec, in their large and base variations, were examined for their ability to capture speech features related to dysphonia. The findings suggest that larger variants of SS models generally outperform smaller ones, with the HuBert and WavLM large models achieving an accuracy of 93.06% and 91.67% in mono-lingual experiments, respectively. Additionally, the study explored cross-lingual capabilities and found that, except for Wav2vec2.0, base variations of SS models exhibited higher accuracies. The highest accuracy achieved in the cross-lingual case was 88.33% by the Wav2vec2.0 model when algorithms were trained on Hungarian samples and tested in the Dutch language. These results highlight the potential of SS models for language-dependent and independent dysphonia detection.