错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Investigation of Video Vision Transformers for Depression Severity Estimation from Facial Video Data

  • Ghazal Bargshady,
  • Roland Goecke

摘要

Recognising depression from facial expressions and movements in video data using machine learning models has gained considerable attention in recent years. Researchers have explored various approaches and techniques to develop models capable of detecting depression-related patterns in facial video data. Recently, Video Vision Transformers have emerged as a powerful deep learning architecture for analysing sequential data, such as video data. While vision transformers have primarily gained attention in computer vision tasks involving images, their application to video analysis tasks, such as the recognition of depression or the estimation of depression severity from facial video data, is an active area of research. In this paper, two different architectures of vision transformers are used to capture spatio-temporal, facial information relevant to estimating the severity of depression and, thus, to provide valuable insights for depression analysis. The models are trained and evaluated on the AVEC2013 and AVEC2014 datasets. The results indicate that the fine-tuned vision transformers can outperform earlier deep learning models in visual depression analysis, achieving a Root Mean Square Error (RMSE) of 5.73 for the vision transformer and 5.39 for the video vision transformers, respectively.