In this work, we propose a novel adaptive key frame mining strategy to guide multi-object tracking, which combines the advantages of short-term and long-term correlation to capture both the structured spatial relationships between different objects as well as the temporal relationships of objects in different frames. We designed the KFE (Key Frame Extraction) module to adaptively perform video segmentation using the reinforcement learning KFE module, which guides the tracker to mine the intrinsic logic of the video to make the correlation effect more significant. In video, mutual occlusion between objects frequently occurs, so the interactions between objects within frames cannot be ignored. Most of the current graph-based work utilizes the features of inter-frame objects for mutual fusion. Based on this, we design an IFF (intra-frame feature fusion) module that passes information between the target and the surrounding objects through GCN to make the target recognizable, thus solving the problem of tracking loss and similar appearance due to object occlusion. Our proposed tracker utilizes both long and short trajectories and takes into account the spatial relationship between objects while leveraging the intrinsic logic of the video clips, thus further improving the performance. It achieves impressive results: 68.6 HOTA, 81.0 IDF1, 66.6 AssA, and 893 IDS on the MOT17 dataset, proving its effectiveness and accuracy.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Spatio-temporal Graph Learning on Adaptive Mined Key Frames for High-Performance Multi-Object Tracking

  • Futian Wang,
  • Fengxiang Liu,
  • Xiao Wang

摘要

In this work, we propose a novel adaptive key frame mining strategy to guide multi-object tracking, which combines the advantages of short-term and long-term correlation to capture both the structured spatial relationships between different objects as well as the temporal relationships of objects in different frames. We designed the KFE (Key Frame Extraction) module to adaptively perform video segmentation using the reinforcement learning KFE module, which guides the tracker to mine the intrinsic logic of the video to make the correlation effect more significant. In video, mutual occlusion between objects frequently occurs, so the interactions between objects within frames cannot be ignored. Most of the current graph-based work utilizes the features of inter-frame objects for mutual fusion. Based on this, we design an IFF (intra-frame feature fusion) module that passes information between the target and the surrounding objects through GCN to make the target recognizable, thus solving the problem of tracking loss and similar appearance due to object occlusion. Our proposed tracker utilizes both long and short trajectories and takes into account the spatial relationship between objects while leveraging the intrinsic logic of the video clips, thus further improving the performance. It achieves impressive results: 68.6 HOTA, 81.0 IDF1, 66.6 AssA, and 893 IDS on the MOT17 dataset, proving its effectiveness and accuracy.