Synthetic Data: Comparing Utility and Risk in Microdata and Tables
摘要
Synthetic data has begun to show potential as an alternative to traditional SDC methods in specific use cases. This development and the increasing research efforts further hint at an emerging role in future privacy protection. However, since data synthesis predominantly happens at microdata level, development of utility and risk metrics is also focused on this domain. Statistical agencies on the other hand limit data publication mostly to aggregates, by selecting various subsets of variables for cross tabulation. We analyze the correlations between microdata and tabular data metrics for assessing utility and risk. Using a large real life data set as an example for data synthesis, we show that certain global metrics may disproportionately represent small subsets of variables, making them an inappropriate estimator for the quality of aggregates. On the other hand, we show strong similarities between certain microdata level risk metrics and risks of group disclosure in aggregated data.