Road Damage Target Detection Using Attention Mechanism and Convolutional Dense Scale Feature Fusion
摘要
Objective: This study focuses on addressing the issue of road damage detection, including the detection of transverse cracks, longitudinal cracks, alligator cracks, potholes, and uneven manhole covers. The goal is to propose an efficient and accurate detection method for instance-level detection tasks in the field of video or image processing, overcoming the limitations of existing technologies, such as the inapplicability of causality to image tasks and over-reliance on multi-stage encoders or decoders. Methods: Using Adaptive Blending technology on five datasets obtained from vehicle-mounted systems, the network architecture was redesigned to include innovative components such as the coordinate attention module and the globally positioned local refinement head. These enhancements improve feature extraction and model inference capabilities, ultimately achieving automatic detection of small targets from a mobile vehicle perspective. Results: On the five target datasets, the algorithm’s average processing time per image is approximately 9.9 ms. The mean Average Precision (mAP) at 0.5 is 91.14%, and the mAP at 0.5:0.95 is 55.46% on the test set. Conclusion: Experimental results show that the algorithm can accurately identify the five types of targets in vehicle-mounted road scenarios while meeting real-time requirements, making it deployable for detecting these five targets on motor vehicles.