错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A transformer-based convolutional local attention (ConvLoA) method for temporal action localization

  • Sainithin Artham,
  • Soharab Hossain Shaikh

摘要

In the realm of temporal localization in videos, our research introduces a novel framework that achieves significant results in event localization in videos. We depart from conventional approaches that rely heavily only on global context encoding for sequence analysis from video. We propose a novel framework that leverages an encoder-decoder mechanism powered by VidSwin to extract global features, which are subsequently combined with the local context. To achieve this, we designed ConvLoA, a convolutional local attention mechanism dedicated to computing contextual focus within localized areas in video frames. ConvLoA extends beyond localization, providing a pathway for generative models to create novel, unseen instructional videos. Extensive experiments on the YouCook2 and ActivityNet datasets were performed. The results affirm that the proposed approach is on par with other state-of-the-art alternatives, validating its competitiveness. This research not only highlights the importance of local context for precise localization but also sets the stage for enhanced video understanding, offering a versatile solution for event localization within videos.