DPETR: a new image-based depth-aware position embedding transformation model for UAV 3D object detection
摘要
Unmanned aerial vehicles (UAVs) drive demand for advanced 3D perception. However, existing approaches face critical limitations: dense Bird’s-Eye View (BEV) methods incur high computational costs and poor long-range performance, while sparse query methods exhibit deficiencies in dynamic detection. To address these, we propose Depth-aware Position Embedding TRansformation (DPETR), a novel 3D detection paradigm. The DPETR integrates three core modules: a hybrid depth estimation network that combines probability buckets with direct regression for precise absolute depth; a 3D adaptive query generation module enhanced through auxiliary 2D proposals; and Geometry Self-Attention module, which leverages depth maps as priors to guide the fusion of RGB and depth features. Experiments on real-world (nuScenes) and simulated UAV (Carla Drone) datasets validate its superiority. The DPETR boosts mAP by 3.2% on nuScenes validation and 8.2% on CDrone test versus baselines. It shows significant improvements in long-range detection accuracy and cross-scenario generalization.