TSAVD-CLIP: One visual encoder learn both temporal and spatial feat
摘要
While significant research efforts have been devoted to Weakly Supervised Video Anomaly Detection (WSVAD), prevailing approaches typically employ a two-stage paradigm: extracting frame-level features via visual encoders followed by temporal modeling on pre-computed features. Although effective, this decoupled framework inherently suffers from information degradation during feature extraction and suboptimal spatiotemporal fusion. To address these limitations, we propose a novel spatiotemporal adaptive framework that integrates temporal, spatial, and fusion adapters into the CLIP image encoder, enabling simultaneous spatiotemporal feature learning at the encoder level. Furthermore, we design a subspace alignment loss to ensure consistent embedding compatibility between the enhanced visual features and the original CLIP text space.Through comprehensive experimental validation on the UCF-Crime benchmark dataset, our framework demonstrates superior efficacy.