<p>In gastroscopy image diagnosis, the physiological structure of the stomach and peristalsis causes many interference and blurred regions in the images, which increases the difficulty of lesion identification. Anchor-based object detection technology is a method that presets anchors with specific scales and aspect ratios in the image to determine the position and category of the target object. It can quickly and accurately identify various objects in complex scenes and provides an important basis for subsequent analysis and decision-making. However, anchor-based object detection techniques in gastroscopy image diagnosis face many issues including limited localization accuracy, high computational cost, and scale sensitivity. Advanced anchor-free object detectors like fully convolutional one-stage object detection (FCOS) were proposed to solve the problems of anchor-based detectors. Unfortunately, for gastroscope image analysis, FCOS-based detectors have three drawbacks: (1) The centrality feature may not work well if the target center does not align with the center of the bounding box. (2) Multi-scale feature fusion ignores the balance of non-adjacent feature map fusion. (3) Feature information is lost before performing classification tasks. To address these issues, we put forward gastric precancerous lesions detector (GPDet), featuring a integrates dual center-ness fusion, a criss-cross-balanced feature pyramid, and fusion of fully connected layers. The dual center-ness fusion combines the semantic and location information of targets, enriching the feature content of center-ness. The criss-cross balanced feature pyramid focuses on the fusion of non-adjacent features, enhancing feature augmentation. The fully connected layer fusion reduces feature information loss in classification tasks, improving GPDet’s feature sensitivity. To assess the effectiveness of GPDet, extensive tests were conducted on a collected gastric precancerous lesions dataset (GPD) along with two public datasets: HyperKvasir and Kvasir-Instrument. The findings clearly demonstrate that GPDet outperforms the existing solutions in terms of performance across these datasets, with its mean average precision (mAP) reaching 55.3%, 67.6%, and 74%, respectively. Notably, when contrasted with FCOS (ResNet101), GPDet exhibits significant.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

GPDet: an anchor-free object detector based on dual center-ness and criss-cross balance for unstructured gastroscopic image data

  • Zhimin Tang,
  • Yuhui Deng,
  • Yi Zhou,
  • Hexian Lu,
  • Lijuan Lu,
  • Junhao Huang,
  • Hong Li,
  • Shun Long

摘要

In gastroscopy image diagnosis, the physiological structure of the stomach and peristalsis causes many interference and blurred regions in the images, which increases the difficulty of lesion identification. Anchor-based object detection technology is a method that presets anchors with specific scales and aspect ratios in the image to determine the position and category of the target object. It can quickly and accurately identify various objects in complex scenes and provides an important basis for subsequent analysis and decision-making. However, anchor-based object detection techniques in gastroscopy image diagnosis face many issues including limited localization accuracy, high computational cost, and scale sensitivity. Advanced anchor-free object detectors like fully convolutional one-stage object detection (FCOS) were proposed to solve the problems of anchor-based detectors. Unfortunately, for gastroscope image analysis, FCOS-based detectors have three drawbacks: (1) The centrality feature may not work well if the target center does not align with the center of the bounding box. (2) Multi-scale feature fusion ignores the balance of non-adjacent feature map fusion. (3) Feature information is lost before performing classification tasks. To address these issues, we put forward gastric precancerous lesions detector (GPDet), featuring a integrates dual center-ness fusion, a criss-cross-balanced feature pyramid, and fusion of fully connected layers. The dual center-ness fusion combines the semantic and location information of targets, enriching the feature content of center-ness. The criss-cross balanced feature pyramid focuses on the fusion of non-adjacent features, enhancing feature augmentation. The fully connected layer fusion reduces feature information loss in classification tasks, improving GPDet’s feature sensitivity. To assess the effectiveness of GPDet, extensive tests were conducted on a collected gastric precancerous lesions dataset (GPD) along with two public datasets: HyperKvasir and Kvasir-Instrument. The findings clearly demonstrate that GPDet outperforms the existing solutions in terms of performance across these datasets, with its mean average precision (mAP) reaching 55.3%, 67.6%, and 74%, respectively. Notably, when contrasted with FCOS (ResNet101), GPDet exhibits significant.