Dynamic Temporal Shift Feature Enhancement for Few-Shot Action Recognition
摘要
Few-shot action recognition aims to accurately predict unseen action categories from limited data. To capture the complex temporal variations within videos, we focus on spatio-temporal modeling through the enhancement of temporal features. Accordingly, we introduce the Dynamic Temporal Shift Feature Enhancement (DTSFE) framework, which incorporates Inter-intra Temporal Channel Mixing (I2TCM) with the Instance-Guided Multi-Scale Module (IMM). The I2TCM module models temporal channel information both within and between frames, effectively capturing dynamic changes. Simultaneously, the IMM module adaptively filters multi-scale spatio-temporal information to enhance the capture of key features. This approach allows the model not only to leverage critical scale features but also to dynamically adapt to variations in sample features, thus improving the temporal modeling of spatio-temporal features and achieving more efficient few-shot action recognition. Experimental results on four few-shot action recognition datasets demonstrate that DTSFE significantly outperforms existing state-of-the-art methods in terms of accuracy and efficiency, underscoring its superior performance and potential for application.