Spatio-temporal Graph Learning on Adaptive Mined Key Frames for High-Performance Multi-Object Tracking
摘要
In this work, we propose a novel adaptive key frame mining strategy to guide multi-object tracking, which combines the advantages of short-term and long-term correlation to capture both the structured spatial relationships between different objects as well as the temporal relationships of objects in different frames. We designed the KFE (Key Frame Extraction) module to adaptively perform video segmentation using the reinforcement learning KFE module, which guides the tracker to mine the intrinsic logic of the video to make the correlation effect more significant. In video, mutual occlusion between objects frequently occurs, so the interactions between objects within frames cannot be ignored. Most of the current graph-based work utilizes the features of inter-frame objects for mutual fusion. Based on this, we design an IFF (intra-frame feature fusion) module that passes information between the target and the surrounding objects through GCN to make the target recognizable, thus solving the problem of tracking loss and similar appearance due to object occlusion. Our proposed tracker utilizes both long and short trajectories and takes into account the spatial relationship between objects while leveraging the intrinsic logic of the video clips, thus further improving the performance. It achieves impressive results: 68.6 HOTA, 81.0 IDF1, 66.6 AssA, and 893 IDS on the MOT17 dataset, proving its effectiveness and accuracy.