Vision-Based Localization and Pose Estimation of Moving Targets with a Wide Field of View
摘要
This paper proposes a “partitioned calibration + cascade detection” framework to address the ill-posedness and error amplification issues in large-field planar-target calibration. The working volume is divided into sub-regions, where multi-view images of a calibration board are collected, optimized, and fused to create a local extrinsic look-up table (LUT) for high-accuracy “near-field” coordinate computation. A two-stage YOLOv10 cascade is then used: the first stage coarsely locates large carriers to narrow the search domain, and the second stage performs fine detection of small targets on cropped sub-images, achieving sub-millimetre pixel extraction and real-time performance.Experiments in a 480 cm × 270 cm industrial space show that the partitioned strategy limits the ranging error across the entire field of view to ≤ 3 cm with a repeatability standard deviation of 2 cm. The mAP50 reaches 0.95 for large objects and 0.89 for small ones, and single-frame inference takes only 2.1 ms. The error increases slowly and linearly with distance, meeting long-range stability requirements. This system breaks the precision bottleneck of single-point calibration in large areas and provides a reproducible engineering paradigm for smart manufacturing and robot navigation.