错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Class-attention video transformer for engagement prediction

  • Xusheng Ai,
  • Victor Sheng,
  • Chunhua Li,
  • Han Yang,
  • Zhiming Cui

摘要

In this paper, we propose the Class Attention in Video Transformer (CavT), an end-to-end method designed to process both long and short variant-length videos for student engagement prediction. CavT introduces a single vector for class embedding and incorporates the Binary-Order Representatives Sampling (BorS) technique to augment the dataset by adding multiple video sequences. Our method outperforms the state-of-the-art with MSE values of 0.0495 on the EmotiW-EP and 0.0377 on the DAiSEE datasets, providing a robust and scalable solution for engagement prediction.