Improving Early Dementia Detection with Advanced Language Models Based on Linguistic Features
摘要
Dementia, a widespread neurodegenerative condition, presents significant challenges in early diagnosis and intervention. This study investigates innovative methods for detecting dementia by analyzing linguistic patterns in speech transcripts through advanced machine learning techniques. Using datasets from DementiaBank, including the Pitt Corpus and ADReSS challenge datasets, we leveraged Large Language Models (LLMs) to assess cognitive and linguistic features. Our methodology encompassed comprehensive data preprocessing, feature extraction from the ‘Cookie Theft’ picture description test, and model fine-tuning. We explored various transfer learning strategies with pre-trained models such as BERT, DistilBERT, RoBERTa, Mistral, and Llama. Our findings indicate that LLMs, particularly when optimized with Low-Rank Adaptation (LoRA) and Quantization (QLoRA), can achieve a dementia detection accuracy of 96%.