The 4D auto-labeling system, with its potential to enhance data annotation efficiency for 3D object detection, has garnered significant attention. However, its adoption has been hampered by the high costs associated with temporally annotated long-sequential training data and limited generalization capabilities across diverse scenarios. In this paper, we hypothesize that a multi-dataset approach can address these challenges and, accordingly, first introduce a Unified training pipeline for multi-dataset 4D Auto-Labeling, namely Uni4DAL. We recognize that this is a challenging task, primarily due to data-level variations and feature-level inconsistencies among various datasets. Motivated by this understanding, we initially propose a series of Data-Level Alignment (DLA) operations to mitigate potential discrepancies between diverse datasets and ensure synchronized training progress across samples from multiple datasets. Furthermore, to address feature-level inconsistencies, we introduce the Mixed Expert Models Voxel Feature Encoding (MoE-VFE) module, which aims to extract both domain-specific and domain-generalizable features. Additionally, we employ a Domain-Adaptive Hard Example Mining (DA-HEM) technique to leverage both data-level and feature-level consistencies, ensuring that the model pays enhanced attention to the challenging samples during multi-dataset training. Finally, comprehensive experiments demonstrate that Uni4DAL significantly improves performance on the nuScenes, Argoverse2, and Waymo datasets, and exhibits greater robustness with insufficient training data.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Uni4DAL: A Unified Baseline for Multi-dataset 4D Auto-Labeling

  • Zhiyuan Yang,
  • Xuekuan Wang,
  • Wei Zhang,
  • Xiao Tan,
  • Jinchen Lu,
  • Jingdong Wang,
  • Errui Ding,
  • Zhihui Lai,
  • Cairong Zhao

摘要

The 4D auto-labeling system, with its potential to enhance data annotation efficiency for 3D object detection, has garnered significant attention. However, its adoption has been hampered by the high costs associated with temporally annotated long-sequential training data and limited generalization capabilities across diverse scenarios. In this paper, we hypothesize that a multi-dataset approach can address these challenges and, accordingly, first introduce a Unified training pipeline for multi-dataset 4D Auto-Labeling, namely Uni4DAL. We recognize that this is a challenging task, primarily due to data-level variations and feature-level inconsistencies among various datasets. Motivated by this understanding, we initially propose a series of Data-Level Alignment (DLA) operations to mitigate potential discrepancies between diverse datasets and ensure synchronized training progress across samples from multiple datasets. Furthermore, to address feature-level inconsistencies, we introduce the Mixed Expert Models Voxel Feature Encoding (MoE-VFE) module, which aims to extract both domain-specific and domain-generalizable features. Additionally, we employ a Domain-Adaptive Hard Example Mining (DA-HEM) technique to leverage both data-level and feature-level consistencies, ensuring that the model pays enhanced attention to the challenging samples during multi-dataset training. Finally, comprehensive experiments demonstrate that Uni4DAL significantly improves performance on the nuScenes, Argoverse2, and Waymo datasets, and exhibits greater robustness with insufficient training data.