In the field of automated fruit harvesting, challenges such as low accuracy and limited equipment performance are commonly encountered. To enhance model performance while ensuring detection accuracy, this study proposes the YOLO-SP model, an improved version of YOLOv11. By introducing a segmentation detection head and a pose estimation detection head, the model is capable of processing a single input image and outputting detection results for two distinct tasks. To further improve model performance, a new branch is incorporated into the feature fusion layer to separate the fused features, addressing the issue of feature loss related to segmentation task for large field objects. Additionally, to facilitate effective integration and convergence of the two tasks, uncertainty loss weights is introduced. Experimental results demonstrate that YOLO-SP achieves commendable performance on the bell pepper dataset, with mean Average Precision (mAP) values of 80.9% for segmentation and 89.4% for pose estimation. Notably, compared to using YOLOv11-Seg and YOLOv11-Pose concurrently, YOLO-SP exhibits a speed improvement of 65.3% and a reduction in parameter count by 22.2%. These findings confirm that the model achieves high accuracy along with significantly enhanced detection speed.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Improved Multi-task Model for Instance Segmentation and Pose Estimation Based on YOLOv11

  • Yan Jiang,
  • Zhitao Dai,
  • Yu Chen,
  • Zhiqiang Chu

摘要

In the field of automated fruit harvesting, challenges such as low accuracy and limited equipment performance are commonly encountered. To enhance model performance while ensuring detection accuracy, this study proposes the YOLO-SP model, an improved version of YOLOv11. By introducing a segmentation detection head and a pose estimation detection head, the model is capable of processing a single input image and outputting detection results for two distinct tasks. To further improve model performance, a new branch is incorporated into the feature fusion layer to separate the fused features, addressing the issue of feature loss related to segmentation task for large field objects. Additionally, to facilitate effective integration and convergence of the two tasks, uncertainty loss weights is introduced. Experimental results demonstrate that YOLO-SP achieves commendable performance on the bell pepper dataset, with mean Average Precision (mAP) values of 80.9% for segmentation and 89.4% for pose estimation. Notably, compared to using YOLOv11-Seg and YOLOv11-Pose concurrently, YOLO-SP exhibits a speed improvement of 65.3% and a reduction in parameter count by 22.2%. These findings confirm that the model achieves high accuracy along with significantly enhanced detection speed.