Exploring Anchor-Free Object Detection Models for Surgical Tool Detection: A Comparative Study of Faster-RCNN, YOLOv4, and CenterNet++
摘要
The evolution of surgical techniques has led to the widespread adoption of Minimally Invasive Surgeries (MIS), offering benefits such as reduced scarring, shorter hospitalization, and less post-operative pain. Despite these advantages, laparoscopic procedures present challenges including limited vision and maneuverability of surgical tools, requiring advanced assistive technologies. This paper presents a comparative study of three state-of-the-art object detection models—Faster-RCNN, YOLOv4, and CenterNet++—for surgical tool detection in laparoscopic videos, using the m2cai16-tool-locations dataset. Our aim is to explore the potential of anchor-free models, exemplified by CenterNet++, and compare them against widely used anchor-based models. The results demonstrate that CenterNet++ achieves competitive performance, with a mAP \(_{50}\) of 0.926, \(mAP_{75}\) of 0.587, and \(mAP_{50:95}\) of 0.556, particularly excelling in higher IoU thresholds and specific surgical tool classes. YOLOv4 demonstrated the best performance in lower IoU thresholds, achieving the highest \(mAP_{50}\) of 0.950. These findings highlight the potential of anchor-free models in the medical imaging domain and suggest promising avenues for further research and application in surgical tool detection.