<p>To address the challenges of cross-domain target association, identity drift, and temporal inconsistencies in multi-camera video surveillance scenarios, this paper proposes a deep learning-based “Skynet-Tracking” video investigation and analysis system. This system unifies target detection, multi-object tracking, pedestrian re-identification, and cross-camera association within a single framework, taking a holistic modeling approach. At the single-camera level, a ReID-enhanced tracking and association mechanism is introduced to improve identity preservation. At the cross-camera level, an open-set identity library is constructed to achieve dynamic identity management and global trajectory fusion. Simultaneously, temporal modeling and time normalization methods are combined to solve the temporal inconsistency problem between multi-source videos. Furthermore, the system employs a distributed architecture to support efficient processing of multi-camera video streams. Experiments were conducted on MOT17, Market-1501, and a self-built multi-camera dataset. Results show that the proposed method outperforms existing methods in terms of MOTA, IDF1, and cross-camera trajectory continuity, significantly reducing the identity switching problem. Visualization results further validate the system's robustness in complex occlusion and viewpoint changing scenarios.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A deep learning-based video forensic and intelligence analysis system for Skynet–Mizong surveillance

  • Danyang Zhao,
  • Yuanyuan Niu

摘要

To address the challenges of cross-domain target association, identity drift, and temporal inconsistencies in multi-camera video surveillance scenarios, this paper proposes a deep learning-based “Skynet-Tracking” video investigation and analysis system. This system unifies target detection, multi-object tracking, pedestrian re-identification, and cross-camera association within a single framework, taking a holistic modeling approach. At the single-camera level, a ReID-enhanced tracking and association mechanism is introduced to improve identity preservation. At the cross-camera level, an open-set identity library is constructed to achieve dynamic identity management and global trajectory fusion. Simultaneously, temporal modeling and time normalization methods are combined to solve the temporal inconsistency problem between multi-source videos. Furthermore, the system employs a distributed architecture to support efficient processing of multi-camera video streams. Experiments were conducted on MOT17, Market-1501, and a self-built multi-camera dataset. Results show that the proposed method outperforms existing methods in terms of MOTA, IDF1, and cross-camera trajectory continuity, significantly reducing the identity switching problem. Visualization results further validate the system's robustness in complex occlusion and viewpoint changing scenarios.