Long-range feature aggregation and occlusion-aware attention for robust autonomous driving detection
摘要
We address the challenge of robust object detection in autonomous driving, where scenes are cluttered with diverse objects at varying scales and frequent occlusions. Building on Deformable DETR, we propose a novel framework that replaces the original deformable attention modules in both the encoder and decoder with Occlusion-Aware Efficient Transformer Attention (OAETA). This enhanced attention mechanism selectively emphasizes visible parts of partially occluded objects, improving detection reliability under real-world driving conditions. Additionally, we introduce Multi-Scale Long-Range Feature Aggregation (MLFA) to fuse multi-level features directly from the backbone network. By capturing long-range dependencies across different spatial scales, MLFA provides richer contextual information for locating small and distant objects in challenging traffic scenarios. Extensive experiments on public autonomous driving benchmarks show that our approach consistently outperforms the Deformable DETR baseline and other state-of-the-art models. Furthermore, ablation studies indicate that the integration of MLFA with OAETA effectively leverages global context, significantly reducing detection errors caused by partial occlusions, including misclassification and localization inaccuracies. These results affirm that combining long-range feature aggregation and occlusion-aware attention notably advances robust object detection capabilities in autonomous driving.