As a data-driven science, machine learning requires vast amounts of training data and computational resources. However, for highly privacy-sensitive data, it is crucial to protect the privacy of the data during both the training and utilization of machine learning models. In this paper, we propose a privacy-preserving machine learning approach using autoencoders and differential privacy mechanisms to safeguard data privacy while minimizing the impact on data availability. Specifically, we augment logistic regression and ResNet18 models with different architectures of autoencoders to perform data encryption? without compromising the machine learning tasks. Additionally, we employ differential privacy mechanisms to introduce gradient perturbations in the encoding part of the autoencoder, enhancing the algorithm’s security and further protecting data privacy. We also design the cosine similarity between the encoded and original data as a metric for evaluating data privacy, considering model performance, privacy budget, and data privacy collectively to balance data availability and privacy. Extensive experiments conducted on MNIST, CIFAR-10, PathMNIST, and BloodMNIST datasets demonstrate that for simple logistic regression models handling easily classifiable datasets, employing simple autoencoder structures can enhance classification accuracy, with significant performance impact after adding differential privacy. For ResNet18, utilizing convolutional autoencoders for data encryption generally has minimal impact on model classification performance and can even improve accuracy in most cases. Adding differential privacy has minor effects on model classification performance. Selecting appropriate model structures and privacy budgets for different usage scenarios can ensure both data availability and privacy.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Encoder-Based Framework for Privacy-Preserving Machine Learning

  • Jiayun Wu,
  • Wei Ren,
  • Xianchao Zhang,
  • Xianghan Zheng

摘要

As a data-driven science, machine learning requires vast amounts of training data and computational resources. However, for highly privacy-sensitive data, it is crucial to protect the privacy of the data during both the training and utilization of machine learning models. In this paper, we propose a privacy-preserving machine learning approach using autoencoders and differential privacy mechanisms to safeguard data privacy while minimizing the impact on data availability. Specifically, we augment logistic regression and ResNet18 models with different architectures of autoencoders to perform data encryption? without compromising the machine learning tasks. Additionally, we employ differential privacy mechanisms to introduce gradient perturbations in the encoding part of the autoencoder, enhancing the algorithm’s security and further protecting data privacy. We also design the cosine similarity between the encoded and original data as a metric for evaluating data privacy, considering model performance, privacy budget, and data privacy collectively to balance data availability and privacy. Extensive experiments conducted on MNIST, CIFAR-10, PathMNIST, and BloodMNIST datasets demonstrate that for simple logistic regression models handling easily classifiable datasets, employing simple autoencoder structures can enhance classification accuracy, with significant performance impact after adding differential privacy. For ResNet18, utilizing convolutional autoencoders for data encryption generally has minimal impact on model classification performance and can even improve accuracy in most cases. Adding differential privacy has minor effects on model classification performance. Selecting appropriate model structures and privacy budgets for different usage scenarios can ensure both data availability and privacy.