LMS-VDR: Integrating Landmarks into Multi-scale Hybrid Net for Video-Based Depression Recognition
摘要
Recent advancements in deep learning have significantly enhanced facial video-based depression recognition. However, existing models encounter critical limitations that hinder their performance. They struggle with spatially localized facial feature extraction, relying on either a single convolution or a convolution operation with a simplistic attention mechanism, leading to inadequate recognition of depression-relevant patterns. Besides, increasing model depth leads to the extraction of ambiguous, abstract features while losing crucial dynamic facial details. To address these issues, we propose LMS-VDR by combining the Multi-Scale Mixed Attention Module (MSMAM) with landmark-based prior knowledge integration. More specifically, MSMAM synergistically merges channel and spatial attention, forming a mixed attention vector block through vector products. It introduces a dense connection mechanism, directly connecting features of each dimension to the final output, thereby enabling multi-scale diversity feature extraction. The integration of landmarks, initially linearly transformed and later combined with temporal feature sequences, enhances dynamic temporal feature extraction using our proposed Cross Multi-head Self-Attention (CMHSA) block based on self-attention. Experiments on AVEC 2013 and AVEC 2014 datasets validate our method’s efficacy, achieving MAE/RMSE of 6.04/7.68 and 5.98/7.59, respectively. Our proposed method offers a promising direction for clinical depression assessment, demonstrating the potential for significant contributions in this critical domain.