LTE-DETR: Long-Tail Scenes Object Detection via Enhanced Feature Representation
摘要
The detection performance of autonomous vehicle object detection algorithms typically degrades significantly in long-tail scenes, potentially leading to serious accidents. Challenges such as complex backgrounds and substantial variations in object morphology further increase detection difficulty. To address these issues, we propose LTE-DETR, a novel object detector designed for long-tail scenes. Specifically, we design a Global-Local Collaborative Feature Adaptive Enhancement Module (GLAEM) to adaptively enhance features extracted by the backbone network, thereby improving the model’s focus on object-relevant features. Additionally, we design a Deformation-Aware Cross-Scale Feature Fusion Module (DA-CFFM) to enhance the model’s adaptability to diverse target morphologies by incorporating deformable feature learning. Compared with the baseline, our method shows superior performance in object detection tasks under long-tail scenes, which validates the effectiveness of LTE-DETR.