Endoscopy is vital for detecting gastrointestinal tract ailments like esophagitis, gastric cancer, and colon cancer at an early stage. Manual segmentation of diseases and polyps in endoscopic images is time-consuming yet Segmenting diseases and polyps manually in endoscopic images requires a significant investment of time. Thus, leveraging deep learning (DL) for endoscopic image segmentation becomes imperative. While Transformer-based models have demonstrated exceptional performance in image analysis tasks by effectively capturing global representations and modeling long-range dependencies, their ability to precisely capture complex boundaries and small objects remains a challenge. To address this, we propose CTIN, a comprehensive Transformer Integration Network. CTIN leverages Transformer’s strengths in extracting main branch segmentation features and introduces the Global Local Fusion Extractor module (GLFE) to facilitate aggregation and expression of both local and global spatial features. This enhancement preserves boundary details and accurately localizes recalibrated objects. Our method demonstrates cutting-edge performance in mDice, mIoU, mPrecision, and mRecall metrics on the CVCClinicDB DD (2020) dataset benchmarks. Additionally, our approach exhibits greater robustness in handling challenging scenarios such as complex boundaries and thin object segmentation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comprehensive Transformer Integration Network (CTIN): Advancing Endoscopic Disease Segmentation with Hybrid Transformer Architecture

  • Jiaming Liang,
  • Mingdu Zhang,
  • Caiyan Tan,
  • Teng Huang,
  • Xi Zhang,
  • Zheng Zhang,
  • Shegan Gao,
  • Qian Sheng,
  • Yan Pang

摘要

Endoscopy is vital for detecting gastrointestinal tract ailments like esophagitis, gastric cancer, and colon cancer at an early stage. Manual segmentation of diseases and polyps in endoscopic images is time-consuming yet Segmenting diseases and polyps manually in endoscopic images requires a significant investment of time. Thus, leveraging deep learning (DL) for endoscopic image segmentation becomes imperative. While Transformer-based models have demonstrated exceptional performance in image analysis tasks by effectively capturing global representations and modeling long-range dependencies, their ability to precisely capture complex boundaries and small objects remains a challenge. To address this, we propose CTIN, a comprehensive Transformer Integration Network. CTIN leverages Transformer’s strengths in extracting main branch segmentation features and introduces the Global Local Fusion Extractor module (GLFE) to facilitate aggregation and expression of both local and global spatial features. This enhancement preserves boundary details and accurately localizes recalibrated objects. Our method demonstrates cutting-edge performance in mDice, mIoU, mPrecision, and mRecall metrics on the CVCClinicDB DD (2020) dataset benchmarks. Additionally, our approach exhibits greater robustness in handling challenging scenarios such as complex boundaries and thin object segmentation.