Anonymizing Personal Information Using Distribution-Based Data Synthesis
摘要
In this study, an approach to the generation and anonymization of a multi-variate dataset is considered, by means of a simulation, which is established on statistical distributions and the integration of business logic in the model. The development of the model emphasizes the need for a multi-criteria approach based on demographic, bank, financial, personal, and individual characteristics to generate an anonymized data set. The need for specific data implies the development of an algorithm based on the distribution of variables and the business logic of the model. It determines the relationships between variables and their distributions.