Endoscopic Surgical Operation and Object Detection Using Custom Architecture Models
摘要
1.2 million Americans have cholecystectomy each year. Surgeons need strong organ visualization and training to prevent injuring other organs. Advances in computer hardware, cameras, etc. could video the doctor's cholecystectomy surgery. Surgical triplets (instrument, verb, target) were annotated on 1 fps videos and 90,489 frames. Each video frame’s dataset includes . Thus, to ease cholecystectomy surgery, surgeons should be provided with a deep learning model that processes real-time surgical video and displays the instrument, verb, target, and instrument on the video monitor. Thus, physicians need not address organ visibility. Five CNN algorithms are compared in this study. Model performance is assessed using accuracy, precision, recall, F1 score, and hamming loss on the validation dataset. Custom architecture CNN, Resnet50, Vgg19, Vgg16, MobileNetv2, YOLO model, and F1 scores were 0.43, 0.75, 0.77, 0.8, 0.63, and 0.24, respectively. This study proposes YOLOv8-based bounding box medical picture categorization. X-rays, CT scans, and MRIs are reliably detected and classified using the suggested procedure. The advanced YOLOv8 model recognizes many items in real time. A convolutional neural network classifies the object of interest in the described approach. The solution outperforms state-of-the-art algorithms with 96.7% classification accuracy on a public medical image dataset.