In this article, we propose FastVim-YOLO, an advanced object detection network aimed at automating fueling processes and reducing labor costs at gas stations through precise detection of vehicle fuel tank caps. Diverse appearances of fuel tank caps and variable environmental conditions, particularly in unstructured settings and under changing lighting, pose significant detection challenges. To overcome these issues, we have developed FastVim-YOLO, a deep learning-based approach that significantly improves detection accuracy and efficiency. This enhanced model combines the vision transformer-based FastViT with the YOLOv8n model and utilizes Mamba model for further optimization. The network incorporates structural reparameterization and introduces the vim module for improved feature extraction, along with the AFPN (Asymptotic Feature Pyramid Network) structure for effective multi-scale feature fusion. To improve the model's performance across a variety of environmental conditions, particularly with respect to lighting, we utilize CycleGAN for image augmentation. This approach creates a robust training dataset that includes virtual nighttime images. Our experimental results demonstrate FastVim-YOLO's superior performance, achieving significant improvements in detection metrics while reducing the model's computational demands. This work presents a promising approach to automating fueling processes, enhancing the efficiency and cost-effectiveness of future gas station operations.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FastVim-YOLO: Lightweight Fuel Tank Cap Detection Network for Refueling Robots

  • Qian Qiao,
  • Jian Guo,
  • Lu Wang,
  • Yu Guo,
  • Kai An

摘要

In this article, we propose FastVim-YOLO, an advanced object detection network aimed at automating fueling processes and reducing labor costs at gas stations through precise detection of vehicle fuel tank caps. Diverse appearances of fuel tank caps and variable environmental conditions, particularly in unstructured settings and under changing lighting, pose significant detection challenges. To overcome these issues, we have developed FastVim-YOLO, a deep learning-based approach that significantly improves detection accuracy and efficiency. This enhanced model combines the vision transformer-based FastViT with the YOLOv8n model and utilizes Mamba model for further optimization. The network incorporates structural reparameterization and introduces the vim module for improved feature extraction, along with the AFPN (Asymptotic Feature Pyramid Network) structure for effective multi-scale feature fusion. To improve the model's performance across a variety of environmental conditions, particularly with respect to lighting, we utilize CycleGAN for image augmentation. This approach creates a robust training dataset that includes virtual nighttime images. Our experimental results demonstrate FastVim-YOLO's superior performance, achieving significant improvements in detection metrics while reducing the model's computational demands. This work presents a promising approach to automating fueling processes, enhancing the efficiency and cost-effectiveness of future gas station operations.