<p>Siamese tracking algorithms usually take convolutional neural networks (CNNs) as feature extractors owing to their capability of extracting deep discriminative features. However, the convolution kernels in CNNs have limited receptive fields, making it difficult to capture global feature dependencies which is important for object detection, especially when the target undergoes large-scale variations or movement. In view of this, we develop a novel network called effective convolution mixed Transformer Siamese network (SiamCMT) for visual tracking, which integrates CNN-based and Transformer-based architectures to capture both local information and long-range dependencies. Specifically, we design a Transformer-based module named lightweight multi-head attention (LWMHA) which can be flexibly embedded into stage-wise CNNs and improve the network’s representation ability. Additionally, we introduce a stage-wise feature aggregation mechanism which integrates features learned from multiple stages. By leveraging both location and semantic information, this mechanism helps the SiamCMT to better locate and find the target. Moreover, to distinguish the contribution of different channels, a channel-wise attention mechanism is introduced to enhance the important channels and suppress the others. Extensive experiments on seven challenging benchmarks, i.e., OTB2015, UAV123, GOT10K, LaSOT, DTB70, UAVTrack112_L, and VOT2018, demonstrate the effectiveness of the proposed algorithm. Specially, the proposed method outperforms the baseline by 3.5% and 3.1% in terms of precision and success rates with a real-time speed of 59.77 FPS on UAV123.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Effective convolution mixed Transformer Siamese network for robust visual tracking

  • Lin Chen,
  • Yungang Liu,
  • Yuan Wang

摘要

Siamese tracking algorithms usually take convolutional neural networks (CNNs) as feature extractors owing to their capability of extracting deep discriminative features. However, the convolution kernels in CNNs have limited receptive fields, making it difficult to capture global feature dependencies which is important for object detection, especially when the target undergoes large-scale variations or movement. In view of this, we develop a novel network called effective convolution mixed Transformer Siamese network (SiamCMT) for visual tracking, which integrates CNN-based and Transformer-based architectures to capture both local information and long-range dependencies. Specifically, we design a Transformer-based module named lightweight multi-head attention (LWMHA) which can be flexibly embedded into stage-wise CNNs and improve the network’s representation ability. Additionally, we introduce a stage-wise feature aggregation mechanism which integrates features learned from multiple stages. By leveraging both location and semantic information, this mechanism helps the SiamCMT to better locate and find the target. Moreover, to distinguish the contribution of different channels, a channel-wise attention mechanism is introduced to enhance the important channels and suppress the others. Extensive experiments on seven challenging benchmarks, i.e., OTB2015, UAV123, GOT10K, LaSOT, DTB70, UAVTrack112_L, and VOT2018, demonstrate the effectiveness of the proposed algorithm. Specially, the proposed method outperforms the baseline by 3.5% and 3.1% in terms of precision and success rates with a real-time speed of 59.77 FPS on UAV123.