<p>This paper introduces Keyframe-Aware Dynamic YOLO (KAD-YOLO), a novel object-detection system designed for real-time surveillance video analysis. The approach leverages keyframe selection, dynamic YOLO, and the Iterative Beluga Whale Optimization-Guided Co-Search (IBWO-HAS) strategy to optimize computational efficiency while maintaining high detection performance. By using a dual-gate mechanism for keyframe selection, the system processes only frames with significant motion or structural changes, reducing unnecessary computations and inference latency. The KAD-YOLO architecture incorporates a dynamic detection head that adapts convolutional weights based on keyframe features, enhancing object localization even under challenging conditions such as occlusions or lighting variations. Additionally, the system employs the Detection-Aware Grad-CAM++ model to provide visual explanations, improving interpretability. Experimental results tested on publicly available surveillance datasets, such as the MOT Challenge, Oxford Town Centre, and UAV123, demonstrate that KAD-YOLO achieves a mean Average Precision (mAP) of 91.3%, precision of 92.8%, recall of 89.5%, and a frames per second (FPS) rate of 52. The model uses only 980&#xa0;MB of memory, outperforming existing YOLO variants (YOLOv5: 85.2% mAP, 45 FPS, 1024&#xa0;MB; YOLOv7: 87.1% mAP, 42 FPS, 1100&#xa0;MB; YOLOv8: 88.0% mAP, 40 FPS, 1150&#xa0;MB). KAD-YOLO also shows superior adaptability to dynamic scenes, including occlusions and varying light conditions. The paper further discusses the potential for KAD-YOLO to be expanded to multi-camera configurations and its future applications in dynamic surveillance environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Keyframe-aware dynamic YOLO with IBWO-guided co-search and detection-aware Grad-CAM++ for real-time surveillance object detection

  • N. M. Saravana Kumar,
  • P. Suresh

摘要

This paper introduces Keyframe-Aware Dynamic YOLO (KAD-YOLO), a novel object-detection system designed for real-time surveillance video analysis. The approach leverages keyframe selection, dynamic YOLO, and the Iterative Beluga Whale Optimization-Guided Co-Search (IBWO-HAS) strategy to optimize computational efficiency while maintaining high detection performance. By using a dual-gate mechanism for keyframe selection, the system processes only frames with significant motion or structural changes, reducing unnecessary computations and inference latency. The KAD-YOLO architecture incorporates a dynamic detection head that adapts convolutional weights based on keyframe features, enhancing object localization even under challenging conditions such as occlusions or lighting variations. Additionally, the system employs the Detection-Aware Grad-CAM++ model to provide visual explanations, improving interpretability. Experimental results tested on publicly available surveillance datasets, such as the MOT Challenge, Oxford Town Centre, and UAV123, demonstrate that KAD-YOLO achieves a mean Average Precision (mAP) of 91.3%, precision of 92.8%, recall of 89.5%, and a frames per second (FPS) rate of 52. The model uses only 980 MB of memory, outperforming existing YOLO variants (YOLOv5: 85.2% mAP, 45 FPS, 1024 MB; YOLOv7: 87.1% mAP, 42 FPS, 1100 MB; YOLOv8: 88.0% mAP, 40 FPS, 1150 MB). KAD-YOLO also shows superior adaptability to dynamic scenes, including occlusions and varying light conditions. The paper further discusses the potential for KAD-YOLO to be expanded to multi-camera configurations and its future applications in dynamic surveillance environments.