Species-Aware Guidance for Animal Action Recognition with Vision-Language Knowledge
摘要
Species diversity is one of the major differences between animal action recognition and human action recognition, resulting in a series of challenges, e.g., action manifestation diversity, concurrent actions, and long-tailed distribution in datasets. As the same action can be manifested significantly differently among animal species due to their physiological differences, it is crucial for models to distinctively learn various visual content under the same label with species-aware perspectives. However, previous works mainly applied single-species recognition methods to animal datasets, without considering species diversity to address animal action recognition. To fill this gap, we propose a novel animal action recognition approach with specific species guidance by exploring pre-trained vision-language knowledge, namely Species-Aware Guidance (SAG). Firstly, we add word-level species semantics to visual embeddings as guidance, leading the model to focus on relevant regions of target animals in subsequent visual understanding. Then, we apply spatiotemporal modeling in both global and local granularity via a two-branch module to obtain a cross-modal video representation. Finally, sentence-level species-aware semantics is fused with action labels as an overall query, guiding the video representation to output the final action label via the decoder. On two widely used public benchmarks of animal action recognition, for both single-label and multi-label scenarios, SAG archives state-of-the-art performance, e.g., Animal Kingdom ( \(\uparrow \) 5.0%), Mammalnet ( \(\uparrow \) 27.0%) compared to existing methods, especially well-alleviating the problem of long-tailed distributions, demonstrating the effectiveness of species guidance under limited data for training.