Federated learning is a notable method to combine efficient data analysis with privacy protection of data, mainly when legal regulations, privacy concerns, or competitive threats restrict data sharing between entities (clients). Instead, clients collaborate to solve the main problem through targeted updates coordinated by some central server. This paper proposes a novel approach combining federated learning, imputation phase, and resampling step to solve the regression problem. The MissForest method is employed during the imputation phase, while the resampling step utilizes an ML algorithm (GAN). The proposed approach is evaluated numerically on various synthetic and practice-oriented datasets. Subsequently, the associated errors related to the statistical properties of the regression models are measured and compared. The results demonstrate that the final model obtained through the proposed method outperforms client-based models in terms of these error metrics.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FLIRT–An Algorithm to Enhance a Regression Model with Federated Learning and GAN-Based Resampling

  • Przemysław Grzegorzewski,
  • Maciej Romaniuk

摘要

Federated learning is a notable method to combine efficient data analysis with privacy protection of data, mainly when legal regulations, privacy concerns, or competitive threats restrict data sharing between entities (clients). Instead, clients collaborate to solve the main problem through targeted updates coordinated by some central server. This paper proposes a novel approach combining federated learning, imputation phase, and resampling step to solve the regression problem. The MissForest method is employed during the imputation phase, while the resampling step utilizes an ML algorithm (GAN). The proposed approach is evaluated numerically on various synthetic and practice-oriented datasets. Subsequently, the associated errors related to the statistical properties of the regression models are measured and compared. The results demonstrate that the final model obtained through the proposed method outperforms client-based models in terms of these error metrics.