Advancements in 3-D object detection: A comprehensive review
摘要
3-D object detection has become essential for autonomous systems, yet the field remains fragmented due to diverse sensor modalities, fusion strategies, and architectural designs. This review aims to unify current approaches by proposing a taxonomy based on fusion granularity, early, mid, and late fusion, and categorizing methods across key architectural families: monocular, LiDAR-only, multi-modal fusion, and transformer-based models. We systematically examine attention mechanisms for contextual and cross-modal modelling, advancements in backbone networks, and solutions for sensor misalignment, calibration issues, and temporal synchronization. Special emphasis is placed on real-world deployment challenges, including occlusion, environmental variability, adverse weather, scalability, and computational efficiency. Methodologically, we survey state-of-the-art models and benchmark their performance using standardized metrics such as mean Average Precision (mAP) and Intersection over Union (IoU) across popular datasets. Results indicate that transformer-based models (e.g., DETR3D, MonoDETR) achieve improved context reasoning and cross-view feature aggregation, outperforming conventional CNN models in Many multi-view and BEV-based tasks. However, they often incur higher computational costs. Fusion-based models demonstrate enhanced robustness to occlusion and sensor discrepancies. Our discussion highlights trade-offs between accuracy, generalization, and real-time inference capabilities, as well as concerns about cost and scalability critical for commercial deployment. By identifying current Limitations and synthesizing recent trends, this review provides a structured foundation for future research in building unified, adaptive, and efficient 3-D detection frameworks.