Prompt tuning, as a parameter-efficient fine-tuning method, plays a crucial role in the fine-tuning of pre-trained models. However, due to the limited expressive power of smaller pre-trained models, the performance of prompt tuning on these smaller models often falls short compared to the larger pre-trained models. To resolve this issue, we propose a knowledge distillation approach that leverages the knowledge of a larger teacher model to enhance the performance of prompt tuning on smaller models. Through analysis and experiments, we first determine that the logit-based distillation method is more suitable for prompt tuning compared to the feature-based method. Building on the commonly used inter-class relationship distillation, we then design and add a new loss function that enables the student model to learn the inter-instance relationships from the teacher model. This expands the information utilized from the teacher model, thereby further enhancing the distillation effect. Experimental results on multiple tasks in the SuperGLUE benchmark indicate that our method significantly enhances the prompt tuning performance of smaller models, even achieving or surpassing the results of larger teacher models in some tasks. Additionally, our method does not alter the structure of the student model, ensuring that the fine-tuned model retains all the advantages of prompt tuning during inference.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Prompt Tuning for Smaller Pretrained Models via Knowledge Distillation

  • Mengyang Yuan,
  • Bo Lang

摘要

Prompt tuning, as a parameter-efficient fine-tuning method, plays a crucial role in the fine-tuning of pre-trained models. However, due to the limited expressive power of smaller pre-trained models, the performance of prompt tuning on these smaller models often falls short compared to the larger pre-trained models. To resolve this issue, we propose a knowledge distillation approach that leverages the knowledge of a larger teacher model to enhance the performance of prompt tuning on smaller models. Through analysis and experiments, we first determine that the logit-based distillation method is more suitable for prompt tuning compared to the feature-based method. Building on the commonly used inter-class relationship distillation, we then design and add a new loss function that enables the student model to learn the inter-instance relationships from the teacher model. This expands the information utilized from the teacher model, thereby further enhancing the distillation effect. Experimental results on multiple tasks in the SuperGLUE benchmark indicate that our method significantly enhances the prompt tuning performance of smaller models, even achieving or surpassing the results of larger teacher models in some tasks. Additionally, our method does not alter the structure of the student model, ensuring that the fine-tuned model retains all the advantages of prompt tuning during inference.