Temporal Action Localization (TAL) is crucial in video understanding, focusing on identifying and timestamping actions within raw video footage. A critical challenge in TAL is processing the rich spatiotemporal details inherent in videos, traditionally addressed through methods adapted from image processing. The Vision Transformer (VIT) model marked a significant evolution, using a self-attention mechanism for enhanced temporal information blending. Despite these advancements, two key issues remain: insufficient extraction of spatial semantic information at lower levels of feature pyramids and inadequate capture of temporal semantic information at higher levels. To address these challenges, we introduce Adaptive Multi-Scale Convolutional Networks with Optimized Attention (AMC-OA). AMC-OA enhances lower-level features within the pyramid using multi-scale convolutional kernels, enriching spatial contextual semantics. Simultaneously, upper-level features are refined with a temporally-focused contextual enhancement network utilizing residual structures for better temporal understanding. To further improve the model’s capability in handling extensive temporal spans, we integrate an advanced multi-head attention mechanism. Empirical results on benchmarks like THUMOS14 and ActivityNet1.3 demonstrate AMC-OA’s superiority in TAL tasks, significantly improving both spatial and temporal information extraction compared to state-of-the-art models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AMC-OA: Adaptive Multi-Scale Convolutional Networks with Optimized Attention for Temporal Action Localization

  • Rui Yuan,
  • Chun Yuan

摘要

Temporal Action Localization (TAL) is crucial in video understanding, focusing on identifying and timestamping actions within raw video footage. A critical challenge in TAL is processing the rich spatiotemporal details inherent in videos, traditionally addressed through methods adapted from image processing. The Vision Transformer (VIT) model marked a significant evolution, using a self-attention mechanism for enhanced temporal information blending. Despite these advancements, two key issues remain: insufficient extraction of spatial semantic information at lower levels of feature pyramids and inadequate capture of temporal semantic information at higher levels. To address these challenges, we introduce Adaptive Multi-Scale Convolutional Networks with Optimized Attention (AMC-OA). AMC-OA enhances lower-level features within the pyramid using multi-scale convolutional kernels, enriching spatial contextual semantics. Simultaneously, upper-level features are refined with a temporally-focused contextual enhancement network utilizing residual structures for better temporal understanding. To further improve the model’s capability in handling extensive temporal spans, we integrate an advanced multi-head attention mechanism. Empirical results on benchmarks like THUMOS14 and ActivityNet1.3 demonstrate AMC-OA’s superiority in TAL tasks, significantly improving both spatial and temporal information extraction compared to state-of-the-art models.