错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Joint Regularization Knowledge Distillation

  • Haifeng Qing,
  • Ning Jiang,
  • Jialiang Tang,
  • Xinlei Huang,
  • Wengqing Wu

摘要

Knowledge distillation is devoted to increasing the similarity between a small student network and an advanced teacher network in order to improve the performance of the student network. However, these methods focus on teacher and student networks that receive supervision from each other independently and do not consider the network as a whole. In this paper, we propose a new knowledge distillation framework called Joint Regularization Knowledge Distillation (JRKD), which aims to reduce network differences through joint training. Specifically, we train teacher and student networks through joint regularization loss to maximize consistency between the two networks. Meanwhile, we develop a confidence-based continuous scheduler method (CBCS), which divides examples into center examples and edge examples based on the example confidence distribution of network output. Prediction differences between networks are reduced when training with a central example. Teacher and student networks will become more similar as a result of joint training. Extensive experimental results on benchmark datasets such as CIFAR-10, CIFAR-100, and Tiny-ImagNet show that JRKD outperforms many advanced distillation methods.