<p>Action anticipation is essential for applications such as autonomous driving, virtual reality, and human-computer interaction. However, existing methods often underexploit the complementary nature of multimodal data, limiting prediction accuracy. To address this issues, we propose a novel multimodal action anticipation network, Cross-Modal Spatiotemporal Fusion with Frequency-Aware (CMSF-FA). We design the FusionComplete-Temporal Feature Aggregation (FCTFA) module for adaptively integrate multimodal features, improving effectiveness. And we propose the Norm-Optimized Vector Attention (NOVA) mechanism and the Frequency-Aware Channel Transformation (FACT) mechanism to enhance the spatiotemporal modeling and frequency-domain information use. We validate the effectiveness of our approach through comparative analysis and ablation studies on two popular benchmark datasets: EPIC-KITCHENS and EGTEA Gaze+.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cross-modal spatiotemporal fusion with frequency-aware for action anticipation

  • Jingru Gui,
  • Yang Lü,
  • Zongnan Ma,
  • Migang Zhang,
  • Fuchun Zhang

摘要

Action anticipation is essential for applications such as autonomous driving, virtual reality, and human-computer interaction. However, existing methods often underexploit the complementary nature of multimodal data, limiting prediction accuracy. To address this issues, we propose a novel multimodal action anticipation network, Cross-Modal Spatiotemporal Fusion with Frequency-Aware (CMSF-FA). We design the FusionComplete-Temporal Feature Aggregation (FCTFA) module for adaptively integrate multimodal features, improving effectiveness. And we propose the Norm-Optimized Vector Attention (NOVA) mechanism and the Frequency-Aware Channel Transformation (FACT) mechanism to enhance the spatiotemporal modeling and frequency-domain information use. We validate the effectiveness of our approach through comparative analysis and ablation studies on two popular benchmark datasets: EPIC-KITCHENS and EGTEA Gaze+.