Zero-Shot Spatio-Temporal Action Detection by Enhancing Context-Relation Capability of Vision-Language Models
摘要
We present a zero-shot spatio-temporal action detection framework that enhances the relational extraction capabilities of vision-language models. Zero-shot spatio-temporal action detection involves identifying a person’s actions in a video and recognizing the time and place of these actions without prior training on those specific actions. Large-scale pre-trained vision-language models like CLIP exhibit zero-shot recognition capabilities for various tasks but struggle with extracting local features and relationships. By explicitly enhancing the extraction of person-context relationships in input videos and improving vision-language feature extraction, our proposed framework performs spatio-temporal action detection. It effectively captures local features and relationships between people and contexts while leveraging the strengths of zero-shot recognition from large-scale vision-language models. The two key components of our framework are person tracking in each input frame while ensuring smooth bounding-box shapes across frames, and the explicit interaction between visual features and language features in the shallow layers of visual feature extraction. We demonstrate the effectiveness of our framework through comprehensive experiments on two well-known action detection datasets, JHMDB and UCF101-24.