<p>In natural scenes, the target pixel ratio in Ochotona curzoniae images is low, and the target features are not prominent, leading to reduced accuracy in feature extraction. Moreover, traditional Ochotona curzoniae object detection models suffer from high memory consumption and limited real-time performance. To address these challenges, we propose an improved object detection model based on YOLOv8. First, we introduce an enhanced channel–spatial attention (ECSA) mechanism that captures cross-channel interactions through one-dimensional convolutions, minimizing information loss due to dimensionality reduction and ensuring a balanced distribution of spatial semantic features. In addition, convolutional kernels of varying sizes are employed to enhance the network’s capability in extracting multi-scale features. Subsequently, a context-guided block (CGB) is integrated into the C2f module of the neck network to improve the localization accuracy and reduce computational complexity. Finally, we propose a context-guided fusion (CGF) module, which calibrates channel weights using a content-guided attention (CGA) mechanism, enabling efficient fusion of shallow and deep features. Experimental results on the Ochotona curzoniae dataset demonstrate that our proposed model achieves a mean average precision (mAP) of 96.4% with a real-time detection speed of 117.1 frames per second (FPS). Compared to the original YOLOv8s model, our approach improves mAP and precision by 2.9% and 8.3%, respectively, while reducing the number of parameters and computational complexity by 8.35% and 5.28%. These results confirm that the proposed model offers both high detection accuracy and real-time performance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An object detection model for Ochotona curzoniae based on the dual-attention mechanism

  • Haiyan Chen,
  • Jie Hao

摘要

In natural scenes, the target pixel ratio in Ochotona curzoniae images is low, and the target features are not prominent, leading to reduced accuracy in feature extraction. Moreover, traditional Ochotona curzoniae object detection models suffer from high memory consumption and limited real-time performance. To address these challenges, we propose an improved object detection model based on YOLOv8. First, we introduce an enhanced channel–spatial attention (ECSA) mechanism that captures cross-channel interactions through one-dimensional convolutions, minimizing information loss due to dimensionality reduction and ensuring a balanced distribution of spatial semantic features. In addition, convolutional kernels of varying sizes are employed to enhance the network’s capability in extracting multi-scale features. Subsequently, a context-guided block (CGB) is integrated into the C2f module of the neck network to improve the localization accuracy and reduce computational complexity. Finally, we propose a context-guided fusion (CGF) module, which calibrates channel weights using a content-guided attention (CGA) mechanism, enabling efficient fusion of shallow and deep features. Experimental results on the Ochotona curzoniae dataset demonstrate that our proposed model achieves a mean average precision (mAP) of 96.4% with a real-time detection speed of 117.1 frames per second (FPS). Compared to the original YOLOv8s model, our approach improves mAP and precision by 2.9% and 8.3%, respectively, while reducing the number of parameters and computational complexity by 8.35% and 5.28%. These results confirm that the proposed model offers both high detection accuracy and real-time performance.