An Investigation of Video Vision Transformers for Depression Severity Estimation from Facial Video Data
摘要
Recognising depression from facial expressions and movements in video data using machine learning models has gained considerable attention in recent years. Researchers have explored various approaches and techniques to develop models capable of detecting depression-related patterns in facial video data. Recently, Video Vision Transformers have emerged as a powerful deep learning architecture for analysing sequential data, such as video data. While vision transformers have primarily gained attention in computer vision tasks involving images, their application to video analysis tasks, such as the recognition of depression or the estimation of depression severity from facial video data, is an active area of research. In this paper, two different architectures of vision transformers are used to capture spatio-temporal, facial information relevant to estimating the severity of depression and, thus, to provide valuable insights for depression analysis. The models are trained and evaluated on the AVEC2013 and AVEC2014 datasets. The results indicate that the fine-tuned vision transformers can outperform earlier deep learning models in visual depression analysis, achieving a Root Mean Square Error (RMSE) of 5.73 for the vision transformer and 5.39 for the video vision transformers, respectively.