Selfish federated learning model trainers may launch data poisoning attacks by introducing specific factors into the data or tampering the sample data, either to obtain a model suitable for themselves or to make the model fail to converge. Detecting participant data poisoning attacks is an essential process in federated learning model training. This paper proposes a gradient similarity-based federated learning data poisoning attack detection scheme. In this scheme, the central server calculates the similarity of the uploaded gradient, and detects anomalies in the uploaded gradients based on historical gradient data, thereby determining whether a user has engaged in data poisoning during the training process. Unlike previous approaches, our method leverages gradient similarity to detect poisoning attacks not only in horizontal federated learning but also in vertical federated learning. In the experiments, we use cosine similarity and Manhattan distance to calculate the similarity of gradient differences uploaded by participants. We consider the participants acting as attacker and analyse the security of this scheme, The experimental results show that the similarity of gradient differences uploaded by honest participants continues to increase as the iterative training progresses, and when the model finally converges, the similarity of gradient differences stabilizes within a small range. When participants engage in local data poisoning, the similarity of gradient differences in uploaded data keeps fluctuating, and the model either fails to converge or the similarity of gradient differences fluctuates within a large range.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Federated Learning Poison Attack Detection Scheme Based on Gradient Similarity

  • Quanyu Zhao,
  • Yuan Zhang,
  • Zhengjun Jing,
  • Yuanjian Zhou,
  • Zexi Xin

摘要

Selfish federated learning model trainers may launch data poisoning attacks by introducing specific factors into the data or tampering the sample data, either to obtain a model suitable for themselves or to make the model fail to converge. Detecting participant data poisoning attacks is an essential process in federated learning model training. This paper proposes a gradient similarity-based federated learning data poisoning attack detection scheme. In this scheme, the central server calculates the similarity of the uploaded gradient, and detects anomalies in the uploaded gradients based on historical gradient data, thereby determining whether a user has engaged in data poisoning during the training process. Unlike previous approaches, our method leverages gradient similarity to detect poisoning attacks not only in horizontal federated learning but also in vertical federated learning. In the experiments, we use cosine similarity and Manhattan distance to calculate the similarity of gradient differences uploaded by participants. We consider the participants acting as attacker and analyse the security of this scheme, The experimental results show that the similarity of gradient differences uploaded by honest participants continues to increase as the iterative training progresses, and when the model finally converges, the similarity of gradient differences stabilizes within a small range. When participants engage in local data poisoning, the similarity of gradient differences in uploaded data keeps fluctuating, and the model either fails to converge or the similarity of gradient differences fluctuates within a large range.