VTA-DepressNet: a dual-level attention fusion for multimodal depression detection
摘要
Depression is a prevalent mental health disorder with significant consequences, making its early and accurate detection essential. Traditional diagnostic methods relying on self-report questionnaires are subjective, underscoring the need for objective, automated approaches. However, unimodal models often fail to capture the full complexity of depressive symptoms. To address this, we propose VTA-DepressNet, a multimodal deep learning architecture that integrates visual, audio, and textual features through an attention-driven fusion mechanism. The model utilizes a Conv-BiLSTM for visual features, a Conv-BiGRU for audio features, and a Transformer encoder with an attention mechanism for textual data. Experimental results on the DAIC-WOZ dataset with fivefold cross-validation demonstrate that VTA-DepressNet achieves an F1-score of 0.83, significantly outperforming unimodal baselines. To validate generalization, the model was further evaluated on the Extended DAIC (E-DAIC) dataset, achieving a competitive F1-score of 0.77. These results not only affirm the effectiveness of combining behavioral and linguistic cues but also highlight the potential of digital mental-health tools for early screening and intervention in both clinical and educational contexts. In practice, VTA-DepressNet could be integrated into telehealth platforms or school-based mental health monitoring systems to support timely and scalable depression assessment.