3D object detection is crucial for autonomous systems, yet its reliance on large-scale annotated LiDAR data creates significant scalability barriers. While semi-supervised learning (SSL) alleviates annotation costs by leveraging unlabeled data, existing SSL-3D detectors face two fundamental limitations: (1) LiDAR’s inherent sparsity and lack of texture features hinder reliable pseudo-label generation, and (2) cross-modal complementary potentials remain underutilized in current single-modality SSL frameworks. In this paper, we propose QuerySS3D, a novel SSL framework that synergizes LiDAR sparse geometric with image dense semantic richness through bidirectional cross-modal interaction. Our approach introduces: (1) A deformable cross-attention mechanism that establishes geometry-semantics co-enhancement by fusing LiDAR queries with image-derived features, and (2) A consistency-driven pseudo-label selection strategy combining Hungarian-aligned geometric matching and confidence-aware tri-clustering. Extensive experiments on KITTI and nuScenes demonstrate that QuerySS3D achieves 70.58% of fully-supervised performance using merely 1% labeled data, reducing annotation costs by 99× while outperforming 3DIoUMatch in recall by 10.2%. Code will be released.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

QuerySS3D: Boosting Semi-supervised 3D Object Detection via Image Query

  • Jiayu Li,
  • Tianzhu Zhang

摘要

3D object detection is crucial for autonomous systems, yet its reliance on large-scale annotated LiDAR data creates significant scalability barriers. While semi-supervised learning (SSL) alleviates annotation costs by leveraging unlabeled data, existing SSL-3D detectors face two fundamental limitations: (1) LiDAR’s inherent sparsity and lack of texture features hinder reliable pseudo-label generation, and (2) cross-modal complementary potentials remain underutilized in current single-modality SSL frameworks. In this paper, we propose QuerySS3D, a novel SSL framework that synergizes LiDAR sparse geometric with image dense semantic richness through bidirectional cross-modal interaction. Our approach introduces: (1) A deformable cross-attention mechanism that establishes geometry-semantics co-enhancement by fusing LiDAR queries with image-derived features, and (2) A consistency-driven pseudo-label selection strategy combining Hungarian-aligned geometric matching and confidence-aware tri-clustering. Extensive experiments on KITTI and nuScenes demonstrate that QuerySS3D achieves 70.58% of fully-supervised performance using merely 1% labeled data, reducing annotation costs by 99× while outperforming 3DIoUMatch in recall by 10.2%. Code will be released.