Advanced Strategies for Privacy Preserving Data Publishing to Improve Multi-class Classification
摘要
Data anonymization has become an essential preprocessing step in many analytics, especially the ones that rely on machine learning. Retrieving external data sources can significantly improve the quality of machine learning models but the sensitive nature often hinders data exchange. On the other hand, anonymization tactics can substantially decrease their utility resulting in inferior machine learning models. This work assesses the quality of anonymized datasets for multi-class classification purposes. A hybrid anonymization pipeline is employed, combining masking and sampling. Deliberate sampling – achieved by balancing the target attribute – not only offers strong and quantifiable privacy guarantees but also improves the utility of released datasets. Our findings are supported by experiments that are executed on three distinct datasets, demonstrating the effectiveness of the approach.