<p>The parameter-efficient fine-tuning (PEFT) has higher efficiency and lower cost for RGB-T tracking, however, it is prone to neglect complex features and modal branch information, compared with the full fine-tuning strategy, which is a major challenge for RGB-T tracking. So, a novel RGB-T tracking method based on multi-modal fusion and adapter fine-tuning (i.e., MFATrack) is proposed in this paper in order to both improve the model's learning ability cross modal features and inherit the training advantages of the PEFT, which including three main components: a bidirectional mask adapter (BMA), a channel interaction feature fusion module (CIFF), and a spatial interaction feature fusion module (SIFF). Based on the PEFT, BMA realizes cross-modal bidirectional feature prompting and establishes dynamic interactions between modalities while reduces training cost. CIFF constructs cross-modal fused template features through channel stitching and a dual aggregation mechanism of multi-modal template features, while SIFF aims to dynamically correlate the fused template features with the multi-modal features in the search area so as to enhance the model's ability to learn complex features and to guide the learning of multi-modal branch information, as effectively addresses tracking challenges in complex scenarios. Experimental results demonstrate that MFATrack achieves PR/SR of 70.3%/56.2% and 87.1%/64.2% on the LasHeR and RGBT234 dataset, as well as PR/SR of 92.3%/75.7% and 85.3%/62.1% on the GTOT and RGBT210 dataset, which validates a good balance between tracking performance and efficiency.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MFATrack: multi-modal fusion tracking network with adapter tunning

  • Ziyu Li,
  • Na You,
  • Tanqing Sun,
  • Mingjia Wang,
  • Xianjun Zhang,
  • Yuping Feng

摘要

The parameter-efficient fine-tuning (PEFT) has higher efficiency and lower cost for RGB-T tracking, however, it is prone to neglect complex features and modal branch information, compared with the full fine-tuning strategy, which is a major challenge for RGB-T tracking. So, a novel RGB-T tracking method based on multi-modal fusion and adapter fine-tuning (i.e., MFATrack) is proposed in this paper in order to both improve the model's learning ability cross modal features and inherit the training advantages of the PEFT, which including three main components: a bidirectional mask adapter (BMA), a channel interaction feature fusion module (CIFF), and a spatial interaction feature fusion module (SIFF). Based on the PEFT, BMA realizes cross-modal bidirectional feature prompting and establishes dynamic interactions between modalities while reduces training cost. CIFF constructs cross-modal fused template features through channel stitching and a dual aggregation mechanism of multi-modal template features, while SIFF aims to dynamically correlate the fused template features with the multi-modal features in the search area so as to enhance the model's ability to learn complex features and to guide the learning of multi-modal branch information, as effectively addresses tracking challenges in complex scenarios. Experimental results demonstrate that MFATrack achieves PR/SR of 70.3%/56.2% and 87.1%/64.2% on the LasHeR and RGBT234 dataset, as well as PR/SR of 92.3%/75.7% and 85.3%/62.1% on the GTOT and RGBT210 dataset, which validates a good balance between tracking performance and efficiency.