Knowledge Distillation (KD) is employed to transfer knowledge from a vast pre-trained teacher network to a reliable and lightweight student network on balanced data, facilitating effective deployment on hardware with limited resources. However, real-world data often exhibits an imbalanced long-tail distribution. The teacher network trained on this data also exhibits imbalanced prediction distributions. The imbalance in both the signals mentioned above significantly impacts student network performance, making training a reliable student network challenging. In this paper, we propose a Virtual Student Distribution Knowledge Distillation (VSD-KD) to alleviate the problem of imbalanced supervision signals. Specifically, we match the teacher imbalance prediction distribution with the virtual student distribution to improve the performance of knowledge transfer, thus mitigating the impact of teacher signal imbalance. Additionally, we adopt methods of class-balanced sampling and class-adaptive data augmentation to reduce the generation of error messages in the head class, thereby alleviating the impact of data signal imbalance. Intensive experiments demonstrate that our method can train a reliable student network on long-tail datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Virtual Student Distribution Knowledge Distillation for Long-Tailed Recognition

  • Haodong Liu,
  • Xinlei Huang,
  • Jialiang Tang,
  • Ning Jiang

摘要

Knowledge Distillation (KD) is employed to transfer knowledge from a vast pre-trained teacher network to a reliable and lightweight student network on balanced data, facilitating effective deployment on hardware with limited resources. However, real-world data often exhibits an imbalanced long-tail distribution. The teacher network trained on this data also exhibits imbalanced prediction distributions. The imbalance in both the signals mentioned above significantly impacts student network performance, making training a reliable student network challenging. In this paper, we propose a Virtual Student Distribution Knowledge Distillation (VSD-KD) to alleviate the problem of imbalanced supervision signals. Specifically, we match the teacher imbalance prediction distribution with the virtual student distribution to improve the performance of knowledge transfer, thus mitigating the impact of teacher signal imbalance. Additionally, we adopt methods of class-balanced sampling and class-adaptive data augmentation to reduce the generation of error messages in the head class, thereby alleviating the impact of data signal imbalance. Intensive experiments demonstrate that our method can train a reliable student network on long-tail datasets.