Optimizing action recognition: a residual convolution with hierarchical and gram matrix based attention mechanisms
摘要
Recognition of human actions based on visual input presents significant challenges due to the diverse ways in which individuals perform identical actions, the temporal fluctuations inherent in these actions, and the variations in perspectives from which they are observed. To overcome the constraints associated with single-modality approaches, researchers have increasingly turned to multimodal visual data fusion techniques. This study introduces an optimized deep learning model, designed for action recognition by integrating various data modalities, including skeletal and inertial data. The model features an innovative 1D convolutional block with residual connections aimed at addressing the vanishing gradient issue, thereby supporting the training of deeper networks and facilitating the extraction of hierarchical features across different temporal scales. Furthermore, a Gram Matrix-based attention mechanism is introduced, which utilizes the similarity among feature vectors over time to create an attention map, thereby improving the model’s capacity to evaluate the significance of long-range temporal dependencies. The model’s efficacy is thoroughly assessed using the publicly accessible UTD-MHAD and CZU-MHAD datasets, revealing superior performance relative to current leading action recognition methodologies with accuracy 93.5% and 94.6% for fused skeleton-inertial on UTD-MHAD and CZU-MHAD datasets respectively. Comprehensive evaluation methods, including confusion matrices, ROC curves, t-SNE visualizations, Grad-CAM heatmaps and attention-kinematic contribution, are employed to validate the model’s robustness and provide insights into its decision-making process. The findings highlight the model’s effectiveness in leveraging multi-modal data fusion and advanced temporal feature extraction, offering a promising approach for improving action recognition tasks.