Enhancing Privacy with Attribute-Centric Anonymization and Synthetic Data Approaches: Comprehensive Reviews and Innovative Solutions
摘要
Privacy preservation in healthcare is essential for safeguarding patient confidentiality, securing sensitive medical data, and ensuring compliance with regulatory standards while enabling meaningful clinical research. To protect individual privacy and facilitate the broad use of personal data for analytics and data mining, anonymization techniques are often employed. Many anonymization methods have been developed over the last few decades, with the most recognized models including k-anonymity, differential privacy, l-diversity, and t-closeness. A greater focus has been placed in recent years on attribute-centric anonymization techniques, which take advantage of the properties of the data being anonymized to improve computational efficiency, privacy and usefulness. Furthermore, synthetic data is also used to protect privacy and meet the growing need for data. This introduced an inclusive review of current literature, which categorizes and analyzes various privacy preservation approaches and synthetic data-based approaches. From this review, the attribute-centric techniques are more robust than the synthetic data-based techniques, because they preserve the original data’s integrity and utility, ensuring more accurate and reliable analysis while protecting privacy. Therefore, in this work, the proposed solution to the research gap is provided by using the attribute-centric technique. The survey includes an examination of more than 65 standard research publications, delving into a variety of technical areas such as different privacy preservation approaches, various attribute-centric and synthetic data-based strategies and performance metrics.