Multi-object tracking (MOT) remains challenging in computer vision due to occlusion and appearance similarity, which hinder tracking accuracy and robustness. This paper introduces TOMTrack, a novel method designed to address these challenges through three core modules. The temporal prediction module builds a temporal information buffer utilizing self-attention to predict the current frame, enhancing spatiotemporal association and improving the accuracy of target position prediction. The feature extraction module employs a block-based strategy for occluded targets, where each block undergoes feature extraction via a CNN before being merged, effectively mitigating occlusion interference and enhancing feature representation. The matching module adopts a two-stage matching strategy to retrieve undetected targets and discover potential targets in the background. These are then integrated to produce the final tracking results and update the temporal information buffer. Experiments on the MOT17 and DanceTrack datasets demonstrate that TOMTrack significantly improves MOT accuracy and robustness, particularly in complex scenarios. The proposed method provides an effective solution to advance MOT.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TOMTrack: Multi-object Tracking with Temporal Info, Occlusion Handling and Object Mining

  • Jiarong Lin,
  • Li Liu

摘要

Multi-object tracking (MOT) remains challenging in computer vision due to occlusion and appearance similarity, which hinder tracking accuracy and robustness. This paper introduces TOMTrack, a novel method designed to address these challenges through three core modules. The temporal prediction module builds a temporal information buffer utilizing self-attention to predict the current frame, enhancing spatiotemporal association and improving the accuracy of target position prediction. The feature extraction module employs a block-based strategy for occluded targets, where each block undergoes feature extraction via a CNN before being merged, effectively mitigating occlusion interference and enhancing feature representation. The matching module adopts a two-stage matching strategy to retrieve undetected targets and discover potential targets in the background. These are then integrated to produce the final tracking results and update the temporal information buffer. Experiments on the MOT17 and DanceTrack datasets demonstrate that TOMTrack significantly improves MOT accuracy and robustness, particularly in complex scenarios. The proposed method provides an effective solution to advance MOT.