Currently, algorithms based on 3D Convolutional Networks have demonstrated remarkable efficacy in the domain of dynamic gesture recognition, exhibiting high levels of accuracy and temporal modeling capabilities. However, these algorithms often involve significant computational cost, with high GFLOPs, which impose stringent hardware requirements and hinder practical applications in the future. The prevailing 3D Convolutional Networks accept a fixed number of video frames as input when processing all categories of gestures. In real-world scenarios, different gestures have varying durations, and the speed of performers’ actions also differs. Therefore, it is important for the network to adapt its input to different gestures, since the GFLOPs of the algorithm is directly related to the number of video frames input to the network. To address this issue, we propose the introduction of the Similarity Guided Sampling (SGS) module to reconstruct the baseline network in dynamic gesture recognition. This module aggregates sliced inputs into groups, enabling the network to adaptively adjust the temporal feature resolution to different gestures. Additionally, we refine the sampling strategy of the module to better preserve crucial information. Experimental results on the EgoGesture Dataset demonstrate that our approach outperforms other methods, striking a balance between high recognition accuracy and reduced computational cost (GFLOPs).

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dynamic Gesture Recognition Using R3D Network with Adaptive Temporal Feature Resolutions

  • Yuren Sun,
  • Hui He,
  • Chaoying Tang,
  • Siyu Huang,
  • Biao Wang,
  • Taiping Jiang

摘要

Currently, algorithms based on 3D Convolutional Networks have demonstrated remarkable efficacy in the domain of dynamic gesture recognition, exhibiting high levels of accuracy and temporal modeling capabilities. However, these algorithms often involve significant computational cost, with high GFLOPs, which impose stringent hardware requirements and hinder practical applications in the future. The prevailing 3D Convolutional Networks accept a fixed number of video frames as input when processing all categories of gestures. In real-world scenarios, different gestures have varying durations, and the speed of performers’ actions also differs. Therefore, it is important for the network to adapt its input to different gestures, since the GFLOPs of the algorithm is directly related to the number of video frames input to the network. To address this issue, we propose the introduction of the Similarity Guided Sampling (SGS) module to reconstruct the baseline network in dynamic gesture recognition. This module aggregates sliced inputs into groups, enabling the network to adaptively adjust the temporal feature resolution to different gestures. Additionally, we refine the sampling strategy of the module to better preserve crucial information. Experimental results on the EgoGesture Dataset demonstrate that our approach outperforms other methods, striking a balance between high recognition accuracy and reduced computational cost (GFLOPs).