Self-discovering temporal anomaly patterns in video anomaly detection via causal representation learning
摘要
Video anomaly detection (VAD) remains limited by environmentally dependent reconstruction losses and annotation scarcity. We propose a causal representation learning approach that formulates anomaly detection as mechanism-invariance learning. Our method encodes video clips into latent motion features and predicts temporal changes through causal inference. Consequently, imposing distributional invariance across temporal windows using maximum mean discrepancy. Unlike existing methods that detect appearance deviations or rely on scene-specific prototypes, our approach uniquely combines autoregressive causal dynamics with windowed residual invariance to enforce law-violation detection. Anomalies are identified when learned motion laws are violated, measured through residual magnitudes and invariance deviations. On benchmarks UCSDped2, Avenue, UCSDped1, and UCF-Crime datasets, our approach achieves 2–5% frame-level area under the curve (AUC) improvements over reconstruction and prediction baselines by reducing false positives under style variations. The framework discovers invariant motion patterns without supervision, addressing scalability challenges in video anomaly detection.