错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DS-SRD: a unified framework for structured representation distillation

  • Yuelong Xia,
  • Jing Yang,
  • Xiaodi Sun,
  • Yungang Zhang

摘要

To improve the representation performance of smaller models, representation distillation has been investigated to transfer structured knowledge from a larger model (teacher) to a smaller model (student). Current work aims to maximize a lower bound on mutual information to transfer global structure knowledge, thus ignoring the local structured semantic transfer from teacher representation. We propose a unified framework for structured representation distillation with the double-student training mechanism, DS-SRD, which focuses on transferring global and local structured, consistent representation between teacher and student. The motivation is that a proficient teacher network has the capability to build a well-structured feature space that accounts for both global and local dependencies. DS-SRD involves guiding the student to mimic more refined structured semantic relations from the teacher, leading to enhanced representation performance. Importantly, this approach is versatile and can be readily extended to various distillation tasks, including supervised representation distillation and self-supervised representation distillation. In addition, we designed a simple CNN-Transformer structure based on DS-SRD to make the CNN encoder attentive via transformer guidance. Extensive experiments were conducted to validate the effectiveness of our method against various supervised and self-supervised representation distillation baselines.