Cross-modal sensors have garnered considerable attention for their ability to seamlessly switch between visible (RGB) and near-infrared (NIR) modalities, allowing them to adapt to varying external lighting conditions. This remarkable feature enables these sensors to provide tracking capabilities under all lighting conditions. However, the distinct imaging mechanisms between these modalities result in significant variations in target appearance, posing challenges for existing tracking algorithms. While a few algorithms have addressed this issue, they often rely on complex module designs and numerous parameters. To tackle these challenges, we propose a Multi-teacher Knowledged Distillation (MKD) algorithm. First, we design two teacher networks to learn modality-specific knowledge from data collected from different modalities. Next, we employ a feature-level knowledge distillation mechanism, using cross-modal data to adaptively transfer knowledge from the teacher networks to the student network. To enhance the stability of the distillation process and accelerate convergence, we introduce a triplet loss function to guide learning. Experimental results on the large-scale cross-modal object tracking dataset (CMOTB) demonstrate the superiority of our method across all evaluation metrics, while avoiding the need for complex module designs.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-teacher Knowledge Distillation with Triplet Loss for Cross-Modal Object Tracking

  • Yi Li,
  • Lei Liu,
  • Mengya Zhang,
  • Chenglong Li

摘要

Cross-modal sensors have garnered considerable attention for their ability to seamlessly switch between visible (RGB) and near-infrared (NIR) modalities, allowing them to adapt to varying external lighting conditions. This remarkable feature enables these sensors to provide tracking capabilities under all lighting conditions. However, the distinct imaging mechanisms between these modalities result in significant variations in target appearance, posing challenges for existing tracking algorithms. While a few algorithms have addressed this issue, they often rely on complex module designs and numerous parameters. To tackle these challenges, we propose a Multi-teacher Knowledged Distillation (MKD) algorithm. First, we design two teacher networks to learn modality-specific knowledge from data collected from different modalities. Next, we employ a feature-level knowledge distillation mechanism, using cross-modal data to adaptively transfer knowledge from the teacher networks to the student network. To enhance the stability of the distillation process and accelerate convergence, we introduce a triplet loss function to guide learning. Experimental results on the large-scale cross-modal object tracking dataset (CMOTB) demonstrate the superiority of our method across all evaluation metrics, while avoiding the need for complex module designs.