Improved YOLO Recognition Algorithm Based on Full-Domain Feature Enhancement and Multi-Scale Grouped Dilated Convolution
摘要
Simultaneous detection of multi-scale targets and full-domain feature extraction are significant challenges in the field of machine vision. Targeting the scenario of identifying potential engineering hazards from highway engineering video surveillance, this paper conducts targeted improvements and enhancements based on the YOLO13 algorithm framework. Firstly, a Multi-scale Grouped Dilated Convolution (MSGDC) module is introduced in both the backbone and neck parts. This module enables efficient multi-scale target feature fusion, achieving simultaneous recognition of numerous multi-scale potential hazards while balancing recognition efficiency and accuracy. Secondly, a Full-domain Transformer (FDT) module is integrated into the backbone part to enhance YOLO’s ability in extracting and expressing full-domain multi-scale features, relying on full-domain features to realize efficient and accurate recognition of multi-scale hazard targets. The improved YOLO13 algorithm is trained and tested on a dataset of potential hazards at highway construction sites. The results show that the improved algorithm exhibits excellent convergence and strong generalization ability, with mAP50 approaching 0.9, precision and recall both near 0.9, AP exceeding 0.95, and the correct classification rate of the confusion matrix close to 1.0. The model demonstrates outstanding detection performance for both large and small targets.