Fusion-Guided Recognition: A Staged Multi-modal Method for Power Line Detection
摘要
Recognizing power lines from aerial images is a critical yet challenging task due to their thin structure and susceptibility to varying visibility conditions. Existing methods, often reliant on single-modality data, struggle in complex scenarios, while multi-modal approaches frequently fail to effectively exploit the complementary strengths of visible (VI) and infrared (IR) imagery. To address these limitations, we propose a novel Multi-modal Staged-Focusing Network (MSF-Net) for robust power line segmentation. Our network consists of a front-end fusion sub-network and a dual-stream recognition backbone. Specifically, the fusion sub-network first produces a rich guidance feature stream from the VI and IR inputs. The core of our framework is the proposed Cross-Stream Enhancement and Aggregation Module (CSEAM), embedded within the recognition backbone. Instead of merely combining features, CSEAM employs a cross-attention mechanism where the VI and IR streams query the guidance stream to dynamically amplify salient features and suppress irrelevant information. Extensive experiments demonstrate that MSF-Net achieves state-of-the-art performance, significantly outperforming existing methods across various challenging scenarios. Our work presents a new paradigm for multi-modal segmentation, showcasing that using fused information as a dynamic guide to enhance, rather than replace, original source features leads to superior robustness and accuracy.