错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Is Transformer Good for Vision-Based Human Action Recognition with Limited Data Source

  • Thanh-Hai Tran,
  • Vuong-Loc Do

摘要

In this day and age, human action recognition (HAR) has become an active research topic as it opens up many applications in video surveillance, healthcare, or entertainment. With the increasing advancement in deep learning and several popular neural networks like convolutional neural networks (CNNs) based HAR has achieved impressive accuracy. Furthermore, transformers, known as one of the powerful tools for Natural Language Processing (NLP) tasks, are slowly replacing CNNs in vision-based tasks because they show the potential to work with huge datasets. In this paper, we conduct a study to investigate whether the transformer model is good for HAR on average or limited data sources. We start to study a state-of-the-art transformer model (e.g. TimeSFormer). After that, we will propose a shortened version of TimeSFormer called LightTimeSFormer by optimizing some parameters in the original architecture to avoid overfitting issues. We finally validate both models on two datasets (HMDB51 and MuWiGes). The experiments confirm and provide recommendations for further research on HAR with CNNs or Transformer models in the future.