SSEM-Net: scene semantic enhancement multimodal network for human action recognition
摘要
Current unimodal approaches for action recognition suffer from inherent limitations: skeleton-based modalities lack adequate spatial interactions, while RGB-based modalities are susceptible to environmental noise. Although existing multimodal methods attempt to integrate information from both modalities to improve performance, they still face challenges such as insufficient cross-modal information fusion and a lack of semantic comprehension. To address these challenges, this paper proposes a Scene Semantic Enhancement Multimodal Network (SSEM-Net) that integrates three modalities: text, skeleton, and RGB. For the text modality branch, a Scene Semantic Enhancement Module is introduced. This module leverages a vision-language large model to extract and integrate contextual scene descriptions, thereby enhancing the model’s comprehension of scene semantics. For the skeletal modality branch, we designed a skeleton-guided RGB frame cropping strategy to suppress background noise and enhance human regions at the input level. Additionally, we propose a dynamic structure-aware graph convolutional network (DSA-GCN) to mitigate over-smoothing of features in graph convolutions, thereby enhancing the discriminative power of skeletal features. During the multimodal fusion stage, a dynamic weight fusion strategy based on the number of valid skeleton points is employed to adaptively balance contributions across modalities. In addition, this paper constructs RTP-Dataset (Rail Transit Passenger Dataset), a dataset specifically designed for detecting anomalous behavior in rail transit scenarios. By studying real-world rail transit cases characterized by complex backgrounds and frequent occlusions, the dataset bridges the gap between laboratory benchmarks and real-world applications. Extensive testing on the self-built dataset and three challenging public datasets demonstrates that our proposed SSEM-Net consistently achieves state-of-the-art performance.