<p>Synthetic data is regarded as a better privacy-preserving alternative to traditionally sanitized data across various applications. However, a recent article challenges this notion, stating that synthetic data does not provide a better trade-off between privacy and utility than traditional anonymization techniques, and that it leads to unpredictable utility loss and highly unpredictable privacy gain. The article also claims to have identified a breach in the differential privacy guarantees provided by PATE-GAN and PrivBayes. Our analysis indicates that when evaluations are conducted in highly specialized and constrained environments, the generalizability of the findings is limited. Moreover, we observed that when key preconditions related to data distributions are not met in experiments, it may lead to spurious observation of violations of the differential privacy guarantee. Subsequently, we performed a comparative privacy-utility analysis using more generalized and unconstrained settings. Our results indicate that although not all synthetic data generation techniques outperformed <i>k</i>-anonymization across every dataset, for each dataset, at least one generator yielded a more favorable privacy-utility trade-off than the <i>k-</i>anonymization method, thereby reaffirming earlier conclusions.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Synthetic data: revisiting the privacy-utility trade-off

  • Fatima Jahan Sarmin,
  • Atiquer Rahman Sarkar,
  • Yang Wang,
  • Noman Mohammed

摘要

Synthetic data is regarded as a better privacy-preserving alternative to traditionally sanitized data across various applications. However, a recent article challenges this notion, stating that synthetic data does not provide a better trade-off between privacy and utility than traditional anonymization techniques, and that it leads to unpredictable utility loss and highly unpredictable privacy gain. The article also claims to have identified a breach in the differential privacy guarantees provided by PATE-GAN and PrivBayes. Our analysis indicates that when evaluations are conducted in highly specialized and constrained environments, the generalizability of the findings is limited. Moreover, we observed that when key preconditions related to data distributions are not met in experiments, it may lead to spurious observation of violations of the differential privacy guarantee. Subsequently, we performed a comparative privacy-utility analysis using more generalized and unconstrained settings. Our results indicate that although not all synthetic data generation techniques outperformed k-anonymization across every dataset, for each dataset, at least one generator yielded a more favorable privacy-utility trade-off than the k-anonymization method, thereby reaffirming earlier conclusions.