OCTAMLLA-UNet: Leveraging Multi-scale Linear Local Attention for Accurate OCTA Retinal Image Segmentation
摘要
Accurate segmentation of vascular structures in Optical Coherence Tomography Angiography (OCTA) images is essential for the effective diagnosis and treatment of retinal diseases. However, the inherent complexity of OCTA images, characterized by significant noise and intricate vessel curvature, presents considerable challenges to existing segmentation methodologies. To address these challenges, we propose OCTA MLLA-UNet, a novel deep learning architecture that embeds Multi-Scale Linear Local Attention (MLLA) module into Unet framework. MLLA-UNet leverages dual-branch geometric encoding, integrating Rotary Position Embedding (RoPE) with Convolutional Position Encoding (CPE) to preserve spatial coherence during rotational transformations, while employing linear attention with depth-wise convolution for global-local dependency modeling. The linear attention core in MLLA module injects rotation-equivariant representations via RoPE-based query/key transformations and local positional encoding. Multi-scale feature stabilization is further enhanced through hierarchical residual connections and stochastic depth regularization. Experimental results demonstrate that MLLA-UNet substantially outperforms state-of-the-art methods, offering a valuable tool for precise retinal vessel segmentation and facilitating early diagnosis and intervention in retinal diseases.