<p>Understanding users’ backchannel agreement is critical for adaptive responsive intelligent systems in education, healthcare, and consultation applications. Traditional landmark-based methods (e.g., those relying on OpenFace landmarks) overlook rich appearance cues, such as appearance changes and subtle muscle movements, which are essential for accurate agreement estimation. To address this, we propose a two-stage, reconstruction-based approach. In the first stage, we fine-tune a Video Masked Autoencoder (VideoMAE) to model facial dynamics. In the second stage, the learned representations are used to estimate backchannel agreement. By segmenting videos into 16-frame chunks, we efficiently capture long-term temporal dynamics for accurate predictions. Experiments on a large public dataset show state-of-the-art performance with an MSE of 0.0576. Ablation studies validate the contributions of each model component. Furthermore, we conduct visualization analyses at both segment and patch levels. These analyses reveal how different temporal segments and facial regions contribute to agreement estimation, providing interpretable behavioral insights. Lastly, our approach balances high accuracy with low inference time, making it suitable for real-time applications. Our proposed framework can extend beyond backchannel agreement to broader affective computing tasks including emotion recognition and engagement detection.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Beyond landmarks: modeling backchannel agreement via reconstruction-based holistic facial motion representations

  • Yuxuan Huang,
  • Eugene Yujun Fu,
  • Peter H. F. Ng

摘要

Understanding users’ backchannel agreement is critical for adaptive responsive intelligent systems in education, healthcare, and consultation applications. Traditional landmark-based methods (e.g., those relying on OpenFace landmarks) overlook rich appearance cues, such as appearance changes and subtle muscle movements, which are essential for accurate agreement estimation. To address this, we propose a two-stage, reconstruction-based approach. In the first stage, we fine-tune a Video Masked Autoencoder (VideoMAE) to model facial dynamics. In the second stage, the learned representations are used to estimate backchannel agreement. By segmenting videos into 16-frame chunks, we efficiently capture long-term temporal dynamics for accurate predictions. Experiments on a large public dataset show state-of-the-art performance with an MSE of 0.0576. Ablation studies validate the contributions of each model component. Furthermore, we conduct visualization analyses at both segment and patch levels. These analyses reveal how different temporal segments and facial regions contribute to agreement estimation, providing interpretable behavioral insights. Lastly, our approach balances high accuracy with low inference time, making it suitable for real-time applications. Our proposed framework can extend beyond backchannel agreement to broader affective computing tasks including emotion recognition and engagement detection.