The introduction of big data and artificial intelligence has led to the generation of immense amounts of data which is extremely hard to control and even harder to protect. Data being a very important aspect of our lives, privacy, and protection of that data has become a major concern. The proposed work is about using privacy-preserving machine learning techniques like differential privacy in the domain of machine learning which is built on random forest algorithms and trained with adult dataset which are completely compatible with the privacy-preserving domain. There have been three test cases in this work, each having their own differences based on the considered constraints of three entities. These three working cases have to be implemented to understand the basic behaviour and learning capabilities of the machine learning model, which would normally have a different result in all different tests. Using the random forest algorithm to train the model and applying differential privacy techniques on the model and the dataset shows us the amount of accuracy that has been compromised to achieve a certain amount of data privacy and protection. The model has been trained on the same dataset three different times while manipulating the noise addition depending on the constraints, trained on random noise, uniform noise, and manual constraints. All the three categories have been noted down including the accuracies to decide which type of data masking would display the required output and is coexisting with the nature of Laplacian noise mechanism.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Differential Privacy-Based Heterogeneous Constraint Learning Model for Data Privacy

  • J. Hyma,
  • M. Ramakrishna Murty,
  • Sirasapalli Joshua Johnson,
  • Praneeth Balabadruni,
  • Sruthi Yenninti

摘要

The introduction of big data and artificial intelligence has led to the generation of immense amounts of data which is extremely hard to control and even harder to protect. Data being a very important aspect of our lives, privacy, and protection of that data has become a major concern. The proposed work is about using privacy-preserving machine learning techniques like differential privacy in the domain of machine learning which is built on random forest algorithms and trained with adult dataset which are completely compatible with the privacy-preserving domain. There have been three test cases in this work, each having their own differences based on the considered constraints of three entities. These three working cases have to be implemented to understand the basic behaviour and learning capabilities of the machine learning model, which would normally have a different result in all different tests. Using the random forest algorithm to train the model and applying differential privacy techniques on the model and the dataset shows us the amount of accuracy that has been compromised to achieve a certain amount of data privacy and protection. The model has been trained on the same dataset three different times while manipulating the noise addition depending on the constraints, trained on random noise, uniform noise, and manual constraints. All the three categories have been noted down including the accuracies to decide which type of data masking would display the required output and is coexisting with the nature of Laplacian noise mechanism.