Uni4DAL: A Unified Baseline for Multi-dataset 4D Auto-Labeling
摘要
The 4D auto-labeling system, with its potential to enhance data annotation efficiency for 3D object detection, has garnered significant attention. However, its adoption has been hampered by the high costs associated with temporally annotated long-sequential training data and limited generalization capabilities across diverse scenarios. In this paper, we hypothesize that a multi-dataset approach can address these challenges and, accordingly, first introduce a Unified training pipeline for multi-dataset 4D Auto-Labeling, namely Uni4DAL. We recognize that this is a challenging task, primarily due to data-level variations and feature-level inconsistencies among various datasets. Motivated by this understanding, we initially propose a series of Data-Level Alignment (DLA) operations to mitigate potential discrepancies between diverse datasets and ensure synchronized training progress across samples from multiple datasets. Furthermore, to address feature-level inconsistencies, we introduce the Mixed Expert Models Voxel Feature Encoding (MoE-VFE) module, which aims to extract both domain-specific and domain-generalizable features. Additionally, we employ a Domain-Adaptive Hard Example Mining (DA-HEM) technique to leverage both data-level and feature-level consistencies, ensuring that the model pays enhanced attention to the challenging samples during multi-dataset training. Finally, comprehensive experiments demonstrate that Uni4DAL significantly improves performance on the nuScenes, Argoverse2, and Waymo datasets, and exhibits greater robustness with insufficient training data.