YoloTransformer-TransDetect: a hybrid model for steel tube defect detection using YOLO and transformer architectures
摘要
Defect detection in industrial environments, especially in steel tube manufacturing, is critical to ensuring product integrity and safety. Traditional methods often have problems with accuracy and efficiency, motivating the exploration of advanced and hybrid techniques such as combining transformer architecture with You Only Look Once (YOLO) object detection. This paper presents a new approach that integrates the convolutional neural networks-Swin Transformer module into the YOLO version 7 (YOLOv7) architecture to enhance feature extraction and error/defect detection on steel tube dataset. Additionally, weighted efficient layer aggregation network, multi-scale channel split, and concatenated convolutional layers are used to further optimize model performance. Together these techniques help optimize the model’s ability to accurately detect and classify errors. Experimental results on the steel tube dataset demonstrate the superiority of the proposed YoloTransformer-TransDetect model compared to YOLOv7, with 5.4% rise in mean average precision, and improvements in precision, and F1 score. This research contributes to the advancement of defect detection hybrid methods promising more accurate and reliable results to ensure steel tube product quality and safety.