<p>The quality of cigarettes is greatly affected by appearance defects. Achieving automatic detection of these defects with high precision and speed has been a critical concern for cigarette factories. To meet the needs of manufacturers in detecting appearance defects in cigarettes, this paper proposes a model based on DETR (DEtection TRansformer) for detecting cigarette appearance defects. The model integrates EfficientViT, SENetV2, and FuNet (Full-scale Feature Fusion Network), called ESF-DETR. First, EfficientViT serves as the backbone feature extraction network, substantially reducing model parameters and enhancing feature extraction efficiency. Second, SENetV2 is introduced at the end of the backbone network to improve feature expression accuracy and global information integration capability. Third, the Full-scale Feature Fusion Network (FuNet) is proposed as the encoder, further reducing model parameters while increasing spatial location and high-level semantic information across each feature layer. The proposed ESF-DETR model achieves a mAP of 96.0% with a parameter count of 10.1M. Compared to the original model, the mAP has increased by 4.4%, while the number of parameters has decreased by 49.8%. Additionally, the detection speed reaches 500 FPS, satisfying cigarette production lines’ accuracy and speed requirements.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ESF-DETR: a real-time and high-precision detection model for cigarette appearance

  • Yingchao Ding,
  • Guowu Yuan,
  • Hao Zhou,
  • Hao Wu

摘要

The quality of cigarettes is greatly affected by appearance defects. Achieving automatic detection of these defects with high precision and speed has been a critical concern for cigarette factories. To meet the needs of manufacturers in detecting appearance defects in cigarettes, this paper proposes a model based on DETR (DEtection TRansformer) for detecting cigarette appearance defects. The model integrates EfficientViT, SENetV2, and FuNet (Full-scale Feature Fusion Network), called ESF-DETR. First, EfficientViT serves as the backbone feature extraction network, substantially reducing model parameters and enhancing feature extraction efficiency. Second, SENetV2 is introduced at the end of the backbone network to improve feature expression accuracy and global information integration capability. Third, the Full-scale Feature Fusion Network (FuNet) is proposed as the encoder, further reducing model parameters while increasing spatial location and high-level semantic information across each feature layer. The proposed ESF-DETR model achieves a mAP of 96.0% with a parameter count of 10.1M. Compared to the original model, the mAP has increased by 4.4%, while the number of parameters has decreased by 49.8%. Additionally, the detection speed reaches 500 FPS, satisfying cigarette production lines’ accuracy and speed requirements.