CMD-CrackNet: A Modular, Interpretable, and Lightweight Approach to Crack Segmentation under Data Scarcity
摘要
Reliable pavement crack segmentation in infrastructure monitoring is hindered by scarce labeled data, complex surface conditions, and the demand for efficient edge deployment. We propose CMD-CrackNet, a lightweight framework combining advances in representation learning and segmentation. The first stage employs AttnCLR, a contrastive self-supervised scheme that integrates multi-head attention into a ResNet encoder, enhancing feature generalization from unlabeled data. The second stage introduces TriDecoderNet, a multi-decoder U-Net with three pathways: (i) a Bayesian decoder for uncertainty-aware predictions, (ii) a tokenization-based decoder for long-range context modeling, and (iii) an adaptive attention decoder for fine spatial refinement. A learnable fusion module integrates these outputs. Efficiency is further improved through a Selective Channel-Spatial Enhancement (SCSE) module, unifying SE and CBAM mechanisms in a compact design. With just 1.04 M parameters, CMD-CrackNet achieves 96.8%-pixel accuracy and an 74.74% Dice score, surpassing deeper baselines. Monte Carlo uncertainty maps and Grad-CAM visualizations enhance interpretability, enabling safe, real-time deployment in critical inspection tasks.