CRCNet: convolutional multi-layer perceptron encoder with attention module for colorectal cancer segmentation
摘要
Colorectal cancer (CRC) segmentation is a difficult task because of the structural complexity of polyps in colonoscopy and endoscopy imaging. This paper proposes CRCNet, a novel deep learning (DL) framework to address these challenges using specifically designed components that improve feature learning and localization precision. CRCNet combines a convolutional multi-layer perceptron (MLP) encoder that preserves both spatial and contextual representations; a synchronized feature boost (SFB) module to maintain fine-grained low-level features usually lost during downsampling. A convolutional block attention module (CBAM) to focus on diagnostically relevant areas using spatial and channel-wise attention. Further, a semantic-local feature aggregation (SLFA) block merges semantic context with localized information to enhance boundary clarity in the decoder. CRCNet is evaluated on four benchmark datasets: CVC-ClinicDB, ETIS, CVC-ColonDB, and CVC-300. It achieves dice coefficient of 0.9519, 0.9139, 0.8677, and 0.9421 intersection over union scores of 0.9083, 0.8438, 0.7692, and 0.8909, respectively. The model also demonstrates exceptional boundary accuracy and lowers mean absolute error (