In response to legislation, companies are now mandated to honor the right to be forgotten by erasing user data. Consequently, it has become imperative to enable data removal in Vertical Federated Learning (VFL), where multiple parties provide private features for model training. In VFL, data removal, i.e., machine unlearning, requires removing specific features across all samples, ensuring privacy in federated learning. To address this challenge, we propose SecureCut, a novel Gradient Boosting Decision Tree (GBDT) framework that effectively enables both instance unlearning and feature unlearning without the need for retraining from scratch. Leveraging a robust GBDT structure, we enable effective data deletion while reducing degradation of model performance. Extensive experimental results on popular datasets demonstrate that our method achieves superior model utility and forgetfulness compared to the state-of-the-art methods. To the best of our knowledge, this is the first work that investigates machine unlearning in VFL scenarios.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SecureCut: Federated Gradient Boosting Decision Trees with Efficient Machine Unlearning

  • Bowen Li,
  • Jian Zhang,
  • Jie Li,
  • Chentao Wu

摘要

In response to legislation, companies are now mandated to honor the right to be forgotten by erasing user data. Consequently, it has become imperative to enable data removal in Vertical Federated Learning (VFL), where multiple parties provide private features for model training. In VFL, data removal, i.e., machine unlearning, requires removing specific features across all samples, ensuring privacy in federated learning. To address this challenge, we propose SecureCut, a novel Gradient Boosting Decision Tree (GBDT) framework that effectively enables both instance unlearning and feature unlearning without the need for retraining from scratch. Leveraging a robust GBDT structure, we enable effective data deletion while reducing degradation of model performance. Extensive experimental results on popular datasets demonstrate that our method achieves superior model utility and forgetfulness compared to the state-of-the-art methods. To the best of our knowledge, this is the first work that investigates machine unlearning in VFL scenarios.