<p>Voice pathology detection (VPD) aims to accurately identify voice impairments by analyzing speech signals. This study proposes models based on deep learning (DL) for binary classification to distinguish between healthy and pathological voices, with a unique focus on the integration of phone posterior probabilities (PPPs as phonetic-based features) alongside <i>Mel</i>-frequency cepstral coefficients (MFCCs as acoustic-based features) for input models. By incorporating PPPs as supplementary information, we investigate the model’s performance across spontaneous, sustained vowel, and read speech datasets, addressing the gap in comparing these speech types for VPD. Our results highlight that PPPs significantly enhance classification accuracy, particularly for read speech and spontaneous speech data types. Using the AVFAD database, we show that the proposed CNN-based model achieves its highest performance on spontaneous speech, with an accuracy of approximately 87% on test data and 93% on validation data. This study emphasizes the impact of PPPs as phonetic-based features in VPD tasks and clarifies which types of speech benefit most from their inclusion, paving the way for more refined models in this field.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluation of phone posterior probabilities for pathology detection in speech data using deep learning models

  • Sahar Farazi,
  • Yasser Shekofteh

摘要

Voice pathology detection (VPD) aims to accurately identify voice impairments by analyzing speech signals. This study proposes models based on deep learning (DL) for binary classification to distinguish between healthy and pathological voices, with a unique focus on the integration of phone posterior probabilities (PPPs as phonetic-based features) alongside Mel-frequency cepstral coefficients (MFCCs as acoustic-based features) for input models. By incorporating PPPs as supplementary information, we investigate the model’s performance across spontaneous, sustained vowel, and read speech datasets, addressing the gap in comparing these speech types for VPD. Our results highlight that PPPs significantly enhance classification accuracy, particularly for read speech and spontaneous speech data types. Using the AVFAD database, we show that the proposed CNN-based model achieves its highest performance on spontaneous speech, with an accuracy of approximately 87% on test data and 93% on validation data. This study emphasizes the impact of PPPs as phonetic-based features in VPD tasks and clarifies which types of speech benefit most from their inclusion, paving the way for more refined models in this field.