Utility Analysis of Differentially Private Anonymized Data Based on Random Sampling
摘要
It is possible to produce differentially private k-anonymized data based on the method of random sampling followed by full-domain generalization for k-anonymization. We previously evaluate the performance of that method, which is implemented as the SafePub algorithm in the ARX anonymization tool. However, since the SafePub algorithm uses the maximum sampling rate that satisfies the requirements for differential privacy, we observe paradoxical results where data utility diminishes as the privacy budget for differential privacy increases. In this paper, we, therefore, conduct preliminary experiments to explore the parameter space for privacy budget and sampling rate by setting the sampling rates explicitly through modifications to the implementation of ARX. Our initial results show the possibility of improving the utility of anonymized data by properly setting the sampling rate below its maximum value.