Towards Automated Screening via Two-Stage Deep Learning: A Pipeline for Classification and Localization of Bleeding from Wireless Capsule Endoscopy Visuals
摘要
This study presents an automated two-stage pipeline utilizing deep learning for gastrointestinal bleeding detection in wireless capsule endoscopy visuals. Through transfer learning, convolutional and YOLOv8 architectures are adapted to demonstrate state-of-the-art performance on an open innovation challenge dataset. The classification model distinguishes between bleeding and non-bleeding frames in a set of 2,618 images with 99.59% accuracy. Subsequently, the object detection model localizes bleeding regions within positive samples, achieving strong average precision of 74.64 and 60.21% at intersection-over-union thresholds of 0.5 and 0.75, respectively. Extensive experimentation validates consistent excellence in generalizing across 30 patient cases exhibiting diverse bleeding manifestations, aided by ample data augmentation. The self-contained solution indicates clinical translation—analyzing vendor-agnostic visuals without parameter tuning needs between datasets. By reliably flagging abnormalities within expansive wireless capsule footage, this system aims to mitigate clinician workload via computer-aided screening, focusing review efforts on critical findings. Our unified recognition model also incorporates localization interpretability via techniques like Eigen CAM to promote practical utility and trust. Progressing automated sensitivity alongside localization precision stands to enhance diagnosis and intervention planning, serving to unlock the abundant visual insights in wireless capsule endoscopy technology towards patient impact.