Single-modality Human Action Recognition (HAR) approaches often fall short in achieving satisfactory performance in recognizing specific actions, attributable to inherent data features. Therefore, we propose a multimodal fusion approach for HAR, rooted in the fusion of skeleton and RGB video. This framework entails training the model separately using both skeleton and RGB video to extract action features intrinsic to each modality. Subsequently, the fusion of skeleton and RGB video is pursued from two key perspectives: classification results and action features. In terms of classification results, we explore and discuss four methods: the confidence-based optimal selection method, the bimodal weighted sum method, and two variations of models reliant on the fusion of classification probability. Regarding action features, we propose a cross-modal fusion training based on action features, operating on both skeleton and RGB video models. This strategy utilizes features from one modality to augment the training of the other, facilitating multi-modality fusion at the feature level. The proposed multimodal fusion strategy is compared to existing methods on two large datasets: NTU RGB+D and NTU RGB+D 120. Experimental results underscore the effectiveness of the propsoed approach.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

RGB+Skeleton: Cross-Modal Fusion Training for Human Action Recognition

  • Ming-Xuan Lin,
  • Wei-Huang Zhang,
  • Hong-Ming Qiu,
  • Hong-Bo Zhang,
  • Miao-Hui Zhang

摘要

Single-modality Human Action Recognition (HAR) approaches often fall short in achieving satisfactory performance in recognizing specific actions, attributable to inherent data features. Therefore, we propose a multimodal fusion approach for HAR, rooted in the fusion of skeleton and RGB video. This framework entails training the model separately using both skeleton and RGB video to extract action features intrinsic to each modality. Subsequently, the fusion of skeleton and RGB video is pursued from two key perspectives: classification results and action features. In terms of classification results, we explore and discuss four methods: the confidence-based optimal selection method, the bimodal weighted sum method, and two variations of models reliant on the fusion of classification probability. Regarding action features, we propose a cross-modal fusion training based on action features, operating on both skeleton and RGB video models. This strategy utilizes features from one modality to augment the training of the other, facilitating multi-modality fusion at the feature level. The proposed multimodal fusion strategy is compared to existing methods on two large datasets: NTU RGB+D and NTU RGB+D 120. Experimental results underscore the effectiveness of the propsoed approach.