BiMambaNet: An efficient crack segmentation method that combines attention mechanism with Mamba
摘要
The semantic segmentation of crack images plays a vital role in geological hazard monitoring and engineering safety assessment. However, due to the complexity of topographic features and the variability of multi-scale textures, Traditional methods struggle to capture fine crack details in complex backgrounds and lack multi-scale structures to effectively address the segmentation requirements of cracks at different scales. To address these issues, this paper proposes a novel Mamba-based U-shaped semantic segmentation network called BiMambaNet. The network employs the BiMamba Local-Global Fusion (BLGF) module as its core module, combined with the Bidirectional Mamba (BiMamba) module to achieve efficient crack spatial structure perception and long-range dependency modeling. The BiMamba module bidirectionally models global context to enhance semantic continuity, while the BLGF module integrates BiMamba module with window attention to achieve a unified representation of local boundary details and global perception. Meanwhile, to enhance the features transmitted through skip connections between the encoder and decoder, a Multi-Scale Attention Fusion (MSAF) module is designed to guide the selection and enhancement of key features through multi-scale receptive fields. Experimental results demonstrate that the proposed network performs well across multiple crack datasets, exhibiting stronger robustness and higher segmentation accuracy, particularly in addressing common challenges in crack images such as large scale variations, complex textures, and blurred boundaries.