<p>Pre-trained language models (PLMs) have emerged as transformative tools in natural language processing. However, their training typically requires massive datasets that are distributed across multiple nodes, posing significant challenges for centralized training paradigms. Federated pre-trained language models (FL-PLMs) offer a promising solution by enabling collaborative model training without raw data. However, they suffer from significant performance degradation due to statistical heterogeneity across clients. Although knowledge distillation has been proposed to address this challenge by aggregating knowledge instead of model parameters, existing approaches remain vulnerable to privacy leakage and potential knowledge degradation during the distillation process. To overcome these limitations, we propose Privacy-preserving Federated Distillation for Pre-trained language models (PFDP), a novel framework designed to optimize both model performance and privacy preservation. PFDP introduces a Local Distillation with Differential Privacy (LoTDP) that safeguards client data privacy. Furthermore, we propose a Global Aggregation via Transfer Learning (GATeL) algorithm that improves the global model’s generalization capabilities while mitigating the catastrophic forgetting phenomenon commonly associated with federated distillation. Extensive evaluation across multiple classification tasks demonstrates the effectiveness of PFDP in improving accuracy while preserving privacy.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

PFDP: privacy-preserving federated distillation method for pretrained language models

  • Chaomeng Chen,
  • Sen Su

摘要

Pre-trained language models (PLMs) have emerged as transformative tools in natural language processing. However, their training typically requires massive datasets that are distributed across multiple nodes, posing significant challenges for centralized training paradigms. Federated pre-trained language models (FL-PLMs) offer a promising solution by enabling collaborative model training without raw data. However, they suffer from significant performance degradation due to statistical heterogeneity across clients. Although knowledge distillation has been proposed to address this challenge by aggregating knowledge instead of model parameters, existing approaches remain vulnerable to privacy leakage and potential knowledge degradation during the distillation process. To overcome these limitations, we propose Privacy-preserving Federated Distillation for Pre-trained language models (PFDP), a novel framework designed to optimize both model performance and privacy preservation. PFDP introduces a Local Distillation with Differential Privacy (LoTDP) that safeguards client data privacy. Furthermore, we propose a Global Aggregation via Transfer Learning (GATeL) algorithm that improves the global model’s generalization capabilities while mitigating the catastrophic forgetting phenomenon commonly associated with federated distillation. Extensive evaluation across multiple classification tasks demonstrates the effectiveness of PFDP in improving accuracy while preserving privacy.