Visual object tracking empowers a wide array of near-earth remote sensing applications based on unmanned aerial vehicles (UAVs). However, the frequent alterations in flight maneuvers and perspectives during UAV tracking present considerable challenges, such as target scale variation and occlusion. Existing template matching approaches often involve complex template updating mechanisms. To mitigate this limitation, we introduce a novel ViTMem network for UAV object tracking. ViTMem comprises three main modules dedicated to precise target localization. Initially, the search image and historical template images features are extracted using Vision Transformer (ViT) encoder. Subsequently, essential template features for target detection in the current frame are dynamically retrieved. Finally, composite features are processed to predict the target localization. Experiments conducted on three demanding datasets demonstrate that ViTMem surpasses state-of-the-art methods, particularly in scenarios with target scale variation and occlusion.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Spatiotemporal Memory Network for UAV Target Tracking

  • Yang Zheng,
  • Yijie Xu,
  • Jimin Liang

摘要

Visual object tracking empowers a wide array of near-earth remote sensing applications based on unmanned aerial vehicles (UAVs). However, the frequent alterations in flight maneuvers and perspectives during UAV tracking present considerable challenges, such as target scale variation and occlusion. Existing template matching approaches often involve complex template updating mechanisms. To mitigate this limitation, we introduce a novel ViTMem network for UAV object tracking. ViTMem comprises three main modules dedicated to precise target localization. Initially, the search image and historical template images features are extracted using Vision Transformer (ViT) encoder. Subsequently, essential template features for target detection in the current frame are dynamically retrieved. Finally, composite features are processed to predict the target localization. Experiments conducted on three demanding datasets demonstrate that ViTMem surpasses state-of-the-art methods, particularly in scenarios with target scale variation and occlusion.