Differentially Private Correlated Attributes Selection for Vertically Partitioned Data Publishing
摘要
This paper delves into the collaborative publishing of vertically partitioned datasets owned by multiple parties in a semi-trusted environment. Currently, solutions for vertically partitioned data publishing are limited and often rely on secure multi-party computation or require training-specific network structures, resulting in high computational overhead and low utility of the synthesized datasets. In response, we proposed a Vertically Partitioned Data Privacy Protection Publishing Scheme (DPCAS). This scheme synthesizes data using selected attribute pairs and their two-way joint distributions, without the need for secure multi-party computation or specific network structures, effectively enhancing data utility and reducing computational overhead. Specifically, each party first performs internal correlated attribute selection and sends the privacy-protected two-dimensional joint distribution of the internally selected attribute pairs and the privacy-protected local datasets to the semi-trusted curator. The curator then integrates ID values based on the received datasets, estimates the low-way joint distributions and mutual information values of cross-party attribute pairs, and selects the cross-party attribute pairs. Finally, the curator iteratively synthesizes the dataset according to the attribute selection results and uses one-way distribution sampling to synthesize missing attributes. Extensive experimental verification demonstrates that DPCAS not only meets differential privacy standards but also significantly enhances the utility of the synthesized data, showcasing promising application prospects.