<p>The intricate biological mechanism governing human vocal production has the remarkable ability to regulate both pitch and volume. However, various internal and external factors can lead to damage in the vocal folds, resulting in alterations in voice quality. These changes not only affect physical functionality but also have an impact on emotional well-being. Hence, the early detection of voice alterations is imperative to facilitate timely intervention and mitigate potential consequences, thereby enhancing the overall quality of life for individuals. In this research on voice disorder detection, an innovative methodology based on Long Short-Term Memory-Domain Adversarial Neural Network is employed. The approach commences with comprehensive data preprocessing, which involves compiling a diverse dataset comprising voice recordings from individuals with both normal and disordered speech. This dataset is sourced from the Arabic Voice Pathology Database. Subsequently, Mel-frequency cepstral coefficients feature extraction is utilized to capture the spectral properties of the voice signals. The crux of the methodology lies in feature classification through LSTM-DANN, where LSTM functions as a feature extractor and DANN facilitates the learning of domain-invariant features. Experimental findings demonstrate an outstanding accuracy rate of 98.89%, thereby affirming the effectiveness of the proposed LSTM-DANN methodology in classifying diverse voice disorders.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Towards Smarter Healthcare: AI-Enhanced Voice Disorder Detection Leveraging LSTM-DANN

  • K. Gowsic,
  • KR Shanmugapriyaa,
  • V. Eswaramoorthy,
  • S. Priyadharshini

摘要

The intricate biological mechanism governing human vocal production has the remarkable ability to regulate both pitch and volume. However, various internal and external factors can lead to damage in the vocal folds, resulting in alterations in voice quality. These changes not only affect physical functionality but also have an impact on emotional well-being. Hence, the early detection of voice alterations is imperative to facilitate timely intervention and mitigate potential consequences, thereby enhancing the overall quality of life for individuals. In this research on voice disorder detection, an innovative methodology based on Long Short-Term Memory-Domain Adversarial Neural Network is employed. The approach commences with comprehensive data preprocessing, which involves compiling a diverse dataset comprising voice recordings from individuals with both normal and disordered speech. This dataset is sourced from the Arabic Voice Pathology Database. Subsequently, Mel-frequency cepstral coefficients feature extraction is utilized to capture the spectral properties of the voice signals. The crux of the methodology lies in feature classification through LSTM-DANN, where LSTM functions as a feature extractor and DANN facilitates the learning of domain-invariant features. Experimental findings demonstrate an outstanding accuracy rate of 98.89%, thereby affirming the effectiveness of the proposed LSTM-DANN methodology in classifying diverse voice disorders.