In many fields, such as healthcare, finance, and scientific research, data sharing and collaboration are critical to achieving better outcomes. However, the sharing of personal data often involves privacy risks, so privacy-preserving techniques are needed to ensure data security and privacy. The superior performance and flexibility of generative models in data representation have led to significant progress in the development of data privacy. This paper proposes a textual data de-privatization scheme based on generative adversarial networks that well combines generative adversarial networks with privacy protection. The generated de-privacy data retains the statistical properties of the original data and effectively removes sensitive information. Moreover, feature sorting and reshaping modules are introduced to enable the generator to better capture the relationships between features, thus improving the quality of synthetic data. In this paper, the utility of synthetic data is evaluated from three aspects, including similarity assessment, privacy assessment, and model utility assessment. Experimental results show that this method achieves a good trade-off in terms of textual data privacy protection and data quality maintenance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Textual Data De-Privatization Scheme Based on Generative Adversarial Networks

  • Yanning Du,
  • Jinnan Xu,
  • Yaling Zhang,
  • Yichuan Wang,
  • Zhoukai Wang

摘要

In many fields, such as healthcare, finance, and scientific research, data sharing and collaboration are critical to achieving better outcomes. However, the sharing of personal data often involves privacy risks, so privacy-preserving techniques are needed to ensure data security and privacy. The superior performance and flexibility of generative models in data representation have led to significant progress in the development of data privacy. This paper proposes a textual data de-privatization scheme based on generative adversarial networks that well combines generative adversarial networks with privacy protection. The generated de-privacy data retains the statistical properties of the original data and effectively removes sensitive information. Moreover, feature sorting and reshaping modules are introduced to enable the generator to better capture the relationships between features, thus improving the quality of synthetic data. In this paper, the utility of synthetic data is evaluated from three aspects, including similarity assessment, privacy assessment, and model utility assessment. Experimental results show that this method achieves a good trade-off in terms of textual data privacy protection and data quality maintenance.