<p>To address the challenges of missed small-target detection, strong background interference, and low detection efficiency in photovoltaic (PV) panel defect inspection, this paper proposes a lightweight improved detection method called YOLO-PVT. Based on the YOLOv8 framework, three enhancement strategies are introduced: (1) the Vision-Aware Attention Convolution (VAConv) module is incorporated, which combines depthwise separable convolution with a local attention mechanism to enhance perception of critical regions; (2) the Transformer-Enhanced Task-Aligned Detection Head (TET-Head) is designed to improve feature collaboration between classification and localization tasks; and (3) the Wise-IoU v2 (WIoUv2) loss function is adopted to optimize regression accuracy for small targets. Experiments conducted on two public PV defect datasets show that YOLO-PVT achieves 92.9% mAP@0.5 and an inference speed of 454.5 FPS with only 3.7M parameters, exhibiting significant improvements in detection accuracy, computational efficiency, and cross-dataset generalization compared to the original YOLOv8 model. Ablation experiments, visual analysis, and cross-dataset evaluations further verify that the model achieves an excellent balance among accuracy, efficiency, and generalization, demonstrating strong potential for practical applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

YOLO-PVT: An Enhanced Algorithm for Photovoltaic Defect Detection in the Presence of Small Targets and Complex Backgrounds

  • Bin Yang,
  • Hui Cai,
  • Changcheng Sun,
  • Shengsheng Qin,
  • Guoqing Miao

摘要

To address the challenges of missed small-target detection, strong background interference, and low detection efficiency in photovoltaic (PV) panel defect inspection, this paper proposes a lightweight improved detection method called YOLO-PVT. Based on the YOLOv8 framework, three enhancement strategies are introduced: (1) the Vision-Aware Attention Convolution (VAConv) module is incorporated, which combines depthwise separable convolution with a local attention mechanism to enhance perception of critical regions; (2) the Transformer-Enhanced Task-Aligned Detection Head (TET-Head) is designed to improve feature collaboration between classification and localization tasks; and (3) the Wise-IoU v2 (WIoUv2) loss function is adopted to optimize regression accuracy for small targets. Experiments conducted on two public PV defect datasets show that YOLO-PVT achieves 92.9% mAP@0.5 and an inference speed of 454.5 FPS with only 3.7M parameters, exhibiting significant improvements in detection accuracy, computational efficiency, and cross-dataset generalization compared to the original YOLOv8 model. Ablation experiments, visual analysis, and cross-dataset evaluations further verify that the model achieves an excellent balance among accuracy, efficiency, and generalization, demonstrating strong potential for practical applications.