3D object detection basing multi-view cameras has gained widespread attention in recent years. Sparse query-based approaches are particularly favored for their low computational cost. However, these methods rely on the transformer architecture, which has limitations in capturing the spatial positional information crucial for accurate 3D object localization. Most current methods still use absolute positional embedding similar to 2D detection tasks, lacking relative positional representations tailored for 3D tasks. To address this, we propose a relative positional embedding method called 3D Rotary Position Embedding (3DRoPE). This method embeds 3D positional information into queries and keys using rotation, then leveraging the properties of rotational calculations to incorporate relative positional information into the subsequent attention weights. Additionally, we introduce a learnable parameter for 3D positional information, allowing the model to adjust to different scales of relative relationships through 3DRoPE. To mitigate the impact of multi-view geometric information on relative position calculations, we incorporate the geometric information into the image’s positional embedding, indirectly enhancing the model’s understanding of different viewpoints as well. 3DRoPE demonstrates superior localization and detection performance compared to previous absolute positional embedding methods on the nuScenes dataset, achieving improvements of 2.4% in mAP and 2.0% in NDS.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Pose-Enhanced 3D Rotary Embedding for Multi-View 3D Object Detection

  • Ke Sheng,
  • Huiying Xu,
  • Xinzhong Zhu

摘要

3D object detection basing multi-view cameras has gained widespread attention in recent years. Sparse query-based approaches are particularly favored for their low computational cost. However, these methods rely on the transformer architecture, which has limitations in capturing the spatial positional information crucial for accurate 3D object localization. Most current methods still use absolute positional embedding similar to 2D detection tasks, lacking relative positional representations tailored for 3D tasks. To address this, we propose a relative positional embedding method called 3D Rotary Position Embedding (3DRoPE). This method embeds 3D positional information into queries and keys using rotation, then leveraging the properties of rotational calculations to incorporate relative positional information into the subsequent attention weights. Additionally, we introduce a learnable parameter for 3D positional information, allowing the model to adjust to different scales of relative relationships through 3DRoPE. To mitigate the impact of multi-view geometric information on relative position calculations, we incorporate the geometric information into the image’s positional embedding, indirectly enhancing the model’s understanding of different viewpoints as well. 3DRoPE demonstrates superior localization and detection performance compared to previous absolute positional embedding methods on the nuScenes dataset, achieving improvements of 2.4% in mAP and 2.0% in NDS.