Hierarchical multi-modal fusion with vision transformers for robust action recognition in infrared-visible videos
摘要
Action recognition in varying illumination and environmental conditions remains challenging, particularly when relying on a single modality. This paper proposes a two-level self-attention fusion framework that integrates visible (RGB) and infrared (IR) video streams using Multi-scale Vision Transformers (MViTs). Each modality’s raw frames and temporal difference frames are processed separately to extract spatial and motion features. Intra-modal fusion is achieved through a multi-head self-attention (MHSA) mechanism, enhancing modality-specific representations. Inter-modal fusion is then performed by applying another MHSA block over concatenated features from the RGB and IR branches, capturing cross-modal dependencies. Experimental results on the Infrared-Visible dataset demonstrate that our dual-stream attention-guided fusion model achieves 96.67% accuracy, significantly outperforming single-modality baselines and conventional fusion techniques. This highlights the effectiveness of our hierarchical fusion strategy and transformer-driven feature learning in multi-modal action recognition.