Multi-cue SORT: integrating weak cues with appearance and motion for multi-object tracking
摘要
The objective of multi-object tracking (MOT) is to accurately detect and track all objects within a continuous sequence while maintaining unique identifiers for each object. Existing research predominantly relies on strong cues from motion and appearance to construct heuristic models, with limited attention given to weak cues arising from object overlap or shape changes as observed by advanced detectors. In this paper, we introduce a novel approach that leverages weak motion cues by adaptively integrating them into high-performance motion versus appearance-based methods. Moreover, by designing weak cue extraction and matching to run independently across targets, our method inherently supports parallelism and GPU acceleration, enabling efficient high-resolution video tracking and showing strong HPC potential. Building upon the appearance-based motion method Deep OC-SORT (in: IEEE International Conference on Image Processing, IEEE, 2023), our approach achieves superior performance on the challenging DanceTrack (in: Proceedings of the IEEE/CVF Conference on Computer 367 Vision and Pattern Recognition, 2022) benchmark, attaining a HOTA score of 61.9. Furthermore, compared to more complex methods, our approach achieves HOTA scores of 65.4 and 64.3 on the MOT17 (Milan in arXiv preprint , 2016) and MOT20 (Dendorfer in arXiv preprint, 2020) benchmarks.