EV2M-YOLOv8: enhancing AI computer vision with EfficientNetV2 and multi-head self-attention for low-complexity agricultural crop, chilli maturity detection
摘要
The high computational complexity of artificial intelligence (AI) computer vision (CV) models has long posed significant challenges for their deployment in physical applications. Elevated computational demands result in higher power consumption and heat generation, reducing efficiency and complicating implementation. These issues are further exacerbated by the rising costs of developing physical applications and the financial burden placed on users to acquire compatible hardware. This paper addresses these challenges by proposing an enhanced model, EV2M-YOLOv8, which aims to effectively reduce computational complexity. The model is based on the YOLOv8 framework but incorporates key modifications: replacing the original backbone with components inspired by EfficientNetV2 convolutional neural network (CNN) modules and integrating Multi-head self-attention (MHSA) mechanisms. The proposed EV2M-YOLOv8, specifically the “s” variant (EV2M-YOLOv8s), was analyzed and trained on datasets focusing on the maturity of agricultural crops, particularly chilli. The results demonstrate the effectiveness of the proposed model. The EV2M-YOLOv8s achieved a high mean average precision (mAP) of 90.3%, reflecting a 0.78% improvement over the original YOLOv8s. Additionally, the computational complexity of the EV2M-YOLOv8s was significantly reduced to 8.7 Giga floating-point operations per second (GFLOPs), representing a 69.37% decrease compared to the original YOLOv8s. These advancements highlight the model’s potential for practical applications, combining enhanced accuracy with lower computational requirements.