Enhancing video pedestrian detection with tracking-information-aided framework and multi-scale feature optimization
摘要
To address high miss rates and low detection accuracy caused by object scale variations and occlusion in video pedestrian detection, this study proposes a General Tracking-Information-Aided Detection Framework (GTIADF). The framework integrates object detection and tracking, leveraging inter-frame temporal information to improve robustness against scale changes and occlusion. GTIADF consists of two main components: (1) an enhanced multi-object tracking algorithm, Deep SORT+, which incorporates advanced modules, camera motion compensation, and DEMA for improved tracking robustness; and (2) an aided detection module that reduces missed detections via validation, interpolation, and re-scoring. The framework is adaptable and can be seamlessly integrated with other detectors to enhance their performance. Based on GTIADF, we propose the D3F-YOLO algorithm, which enhances the feature extraction module by using deformable convolution for better detection at varying scales. Additionally, a focusing diffusion pyramid network is introduced to improve multiscale feature representation, and the loss function is optimized to boost accuracy. Experiments on the Caltech dataset show that the proposed method achieves a mean mean precision (mAP) of 65. 7% and a miss rate of 32.4%, confirming its effectiveness in challenging scenarios.