In practical classification problems, the issue of sample imbalance is a pervasive challenge. The advent of federated learning, which involves the sharing of models among multiple participants without sharing data, has further complicated the handling of sample imbalance. This complexity is particularly pronounced in vertical federated learning, where a high degree of overlap in participant samples is required. The process of aligning samples while preserving privacy may result in a significant reduction in the available data samples, exacerbating the pre-existing imbalance issue. In the context of ensuring data privacy, we propose a secure privacy-preserving SMOTE (SP2-SMOTE) sampling method. It extends traditional SMOTE by allowing parties to independently generate synthetic samples without exposing the data, while effectively preventing unauthorized label inference through minority-class nearest neighbor interpolation. The evaluation of the imbalanced KEEL dataset, divided into two participants based on sample feature importance, demonstrates that SP2-SMOTE significantly improves the classification performance of vertical federated learning. These advances are validated by a series of metrics. This work offers a robust solution to the challenge of imbalanced data in vertical federated learning, rigorously preserving privacy for practical applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Secure Privacy-Preserving SMOTE for Vertical Federated Learning

  • Wenyou Du,
  • Haihang Wang,
  • Jiaming Shen,
  • Guanglei Meng,
  • Yuming Guo,
  • Wei Zhou

摘要

In practical classification problems, the issue of sample imbalance is a pervasive challenge. The advent of federated learning, which involves the sharing of models among multiple participants without sharing data, has further complicated the handling of sample imbalance. This complexity is particularly pronounced in vertical federated learning, where a high degree of overlap in participant samples is required. The process of aligning samples while preserving privacy may result in a significant reduction in the available data samples, exacerbating the pre-existing imbalance issue. In the context of ensuring data privacy, we propose a secure privacy-preserving SMOTE (SP2-SMOTE) sampling method. It extends traditional SMOTE by allowing parties to independently generate synthetic samples without exposing the data, while effectively preventing unauthorized label inference through minority-class nearest neighbor interpolation. The evaluation of the imbalanced KEEL dataset, divided into two participants based on sample feature importance, demonstrates that SP2-SMOTE significantly improves the classification performance of vertical federated learning. These advances are validated by a series of metrics. This work offers a robust solution to the challenge of imbalanced data in vertical federated learning, rigorously preserving privacy for practical applications.