While multimodal fusion has shown promise in video anomaly detection, reliance on audio inputs remains impractical due to hardware limitations and privacy concerns. This paper focuses exclusively on the visual domain, using RGB and optical flow to achieve efficient and accurate detection. We propose a lightweight transformer-based memory mechanism that enhances feature fusion through a knowledge marker weighting scheme. To improve generalization, we introduce a Fusion Feature Enhancement Module (FFEM), which employs random masking, multi-scale convolutions, and attention-based fusion to suppress overfitting and enrich spatial representations. Additionally, we develop Minimum Recently Used Memory (M-LRU), a memory optimization strategy that enables efficient long-duration anomaly detection. Experiments on UCF-Crime and XD-Violence demonstrate the effectiveness and strong cross-dataset generalization of our method, all within a compact model design.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Weakly Supervised Video Anomaly Detection with Lightweight Knowledge Token Fusion and Memory Optimization Strategy

  • Fangxin Hou,
  • Qingxuan Shi,
  • Fang Yang,
  • Sijia Fu,
  • Junhui Li

摘要

While multimodal fusion has shown promise in video anomaly detection, reliance on audio inputs remains impractical due to hardware limitations and privacy concerns. This paper focuses exclusively on the visual domain, using RGB and optical flow to achieve efficient and accurate detection. We propose a lightweight transformer-based memory mechanism that enhances feature fusion through a knowledge marker weighting scheme. To improve generalization, we introduce a Fusion Feature Enhancement Module (FFEM), which employs random masking, multi-scale convolutions, and attention-based fusion to suppress overfitting and enrich spatial representations. Additionally, we develop Minimum Recently Used Memory (M-LRU), a memory optimization strategy that enables efficient long-duration anomaly detection. Experiments on UCF-Crime and XD-Violence demonstrate the effectiveness and strong cross-dataset generalization of our method, all within a compact model design.