Semantic segmentation of marine animal image by U-Net based on multi-cognitive visual adapter and dual-attention fusion mechanism
摘要
Segmenting transparent organisms in underwater environments presents unique challenges due to the optical similarity in refractive index between their body surfaces and seawater, resulting in blurred contours that impede precise delineation. To advance research in underwater transparent segmentation, we first construct and open-source the first large-scale dataset for underwater transparent organisms, TransGlassAqua. This dataset comprises 8,084 annotated images across seven distinct underwater transparent species, providing crucial data support and benchmarks for transparent biological segmentation. Second, to thoroughly validate TransGlassAqua’s effectiveness and address the aforementioned segmentation difficulties, we propose a simple yet efficient segmentation network. It integrates a Multi-cognitive Visual Adapter built upon the Segment Anything Model 2 backbone to enhance multi-scale perception and enable efficient fine-tuning. A Dynamic Dual Attention Fusion module is designed to achieve adaptive, semantic fusion of multi-scale features through hierarchical channel-spatial dual attention and dynamic feature recalibration. Furthermore, multi-scale linear attention blocks are introduced to strengthen contextual feature extraction. Extensive evaluations on TransGlassAqua and four other public datasets validate its merits: DM-UNet achieves state-of-the-art performance on marine animal segmentation tasks. Notably, on our TransGlassAqua dataset, the algorithm attains a mIoU of 88.4%, surpassing current mainstream algorithms by a significant margin of 1.5%. This robustly demonstrates the dataset’s effectiveness in driving transparent object segmentation research. These results offer new insights for underwater transparent segmentation.