Knowledge distillation aims to transfer the knowledge of the pre-trained teacher model to the student model. Many existing approaches align the probabilistic prediction scores of the student model directly with those of the teacher model (i.e., soft labels) without taking into account the capacity of the teacher model. Soft labels generated by strong and weak teachers can respectively result in information loss in non-target class and misguidance in target class. Moreover, traditional alignment methods often lead to poor generalization due to the absence of explicit encouragement for discriminative feature learning. In this paper, we propose an efficient method named Instance-level Scaling and Dynamic Margin-alignment knowledge distillation (ISDM). To balance the impact of the target class and non-target class, we utilize an entropy regularization loss to scale the target class of the teacher model at instance-level. Besides, to enhance the discriminative ability of the student, we incorporate dynamic margin-alignment between the student and teacher models. Our method optimizes the soft labels and improves the discriminative ability of the student model. Experimental results on the CIFAR-100 and ImageNet-1k datasets demonstrate that our method achieves state-of-the-art performance across all teacher-student pairs at a lower cost.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Instance-Level Scaling and Dynamic Margin-Alignment Knowledge Distillation

  • Chuan Li,
  • Xiao Teng,
  • Yan Ding,
  • Changjian Wang,
  • Zheng Qin,
  • Long Lan,
  • Jing Zhang

摘要

Knowledge distillation aims to transfer the knowledge of the pre-trained teacher model to the student model. Many existing approaches align the probabilistic prediction scores of the student model directly with those of the teacher model (i.e., soft labels) without taking into account the capacity of the teacher model. Soft labels generated by strong and weak teachers can respectively result in information loss in non-target class and misguidance in target class. Moreover, traditional alignment methods often lead to poor generalization due to the absence of explicit encouragement for discriminative feature learning. In this paper, we propose an efficient method named Instance-level Scaling and Dynamic Margin-alignment knowledge distillation (ISDM). To balance the impact of the target class and non-target class, we utilize an entropy regularization loss to scale the target class of the teacher model at instance-level. Besides, to enhance the discriminative ability of the student, we incorporate dynamic margin-alignment between the student and teacher models. Our method optimizes the soft labels and improves the discriminative ability of the student model. Experimental results on the CIFAR-100 and ImageNet-1k datasets demonstrate that our method achieves state-of-the-art performance across all teacher-student pairs at a lower cost.