Cross-modal spatiotemporal fusion with frequency-aware for action anticipation
摘要
Action anticipation is essential for applications such as autonomous driving, virtual reality, and human-computer interaction. However, existing methods often underexploit the complementary nature of multimodal data, limiting prediction accuracy. To address this issues, we propose a novel multimodal action anticipation network, Cross-Modal Spatiotemporal Fusion with Frequency-Aware (CMSF-FA). We design the FusionComplete-Temporal Feature Aggregation (FCTFA) module for adaptively integrate multimodal features, improving effectiveness. And we propose the Norm-Optimized Vector Attention (NOVA) mechanism and the Frequency-Aware Channel Transformation (FACT) mechanism to enhance the spatiotemporal modeling and frequency-domain information use. We validate the effectiveness of our approach through comparative analysis and ablation studies on two popular benchmark datasets: EPIC-KITCHENS and EGTEA Gaze+.