SFMB-YOLO: A Deeply Customized Multi-Level Feature Extraction Architecture for Rail Transit Obstacle Detection
摘要
Robust forward obstacle detection in rail transit is fundamentally challenged by the difficulty of accurately and efficiently discerning multi-scale obstacles with sufficient feature clarity from complex, high-resolution track imagery. To address this issue, we introduces SFMB-YOLO, a deeply customized multi-level feature extraction architecture for robust and efficient rail transit obstacle detection. SFMB-YOLO innovatively integrates three core modules: the StarFusion Operation (SFO) block, which enhances non-linear feature aggregation and representation learning without significantly increasing computational complexity; the Mona + block, inspired by human visual cognition, which employs multi-scale depth-wise convolutions for improved perception of variably-sized obstacles; and the BMFormer block, which utilizes bi-level routing attention and a Mixture of Experts (MoE) for efficient long-range contextual modeling. Comprehensive experiments conducted on our custom Rail2D dataset demonstrate that SFMB-YOLO achieves superior performance compared to state-of-the-art methods. This work significantly contributes to advancing forward obstacle detection capabilities in rail transit by providing a lightweight yet powerful solution that effectively balances detection accuracy, robustness to diverse environmental conditions, and real-time processing efficiency.