GovSynBayes: release of synthetic government microdata from multisources via Bayesian networks
摘要
Government agencies respond to policies on open government data and develop the innovation of services by releasing structured microdata. Before release, privacy protection is necessary to mitigate privacy disclosure risks. Synthetic Data Generation technique has attracted more and more attention in the field of privacy protection. The limitations of synthesizing comprehensive data are imposed owing to the nature of isolated government microdata. To release synthetic microdata from multiple government departments while keeping each department’s control over its local data, a framework for releasing privacy-preserving synthetic data is established, and a Bayesian network-based approach GovSynBayes is proposed. Firstly, the count histograms of multidimensional attributes from multi-department data sources are generated through federated queries. Secondly, the histograms are utilized to construct differentially private Bayesian networks. Finally, the synthesized microdata is sampled and released via the generated private Bayesian networks. It is validated that the exponential mechanism is superior for private network learning compared to the Laplace mechanism. A privacy budget allocation algorithm named consistent signal-to-noise for private distribution learning is proposed, which efficiently reduces the average variation distance between synthetic data and original data. GovSynBayes has been experimentally evaluated on four real government micro-datasets. Compared with the currently popular private synthesizers, GovSynBayes has better computational efficiency and generates synthetic dataset with better attribute correlation and machine learning utility.