A Stacked Deep Learning Model for Accurate Dementia Detection Using Transcript Data
摘要
This study uses speech information from the DementiaBank Pitt Corpus to investigate deep learning models for the early identification of dementia. Classifying patients as dementia-positive (AD+) or negative (AD-) is the goal of the models, which examine transcripts of verbal exchanges between patients and therapists. Several architectures were used, such as Transformer, CNN, CNN + BiLSTM, Stacked Deep Dense Neural Network (SDDNN), and Attention-based LSTM. Important measures like accuracy, precision, recall, F1 score, specificity, and AUC were used to assess each model. Pre-trained GloVe embedding models consistently outperformed randomly initialized embedding models; the greatest accuracy (94.00%) and AUC (0.9350) were achieved by the Attention-based LSTM model. The research results demonstrate how well hybrid deep learning models—particularly those that incorporate attention mechanisms—capture complex linguistic features, providing a scalable and effortless means of diagnosing dementia at an early stage.