An Automated Quasi-Identification (QID) for Re-identification
摘要
The rapid growth of data generation and technology has led to concerns about security and privacy. Personal and sensitive information can be disclosed by unauthorized individuals, necessitating the use of privacy-preserving techniques. In this study, an automated quasi-identification (QID) approach for privacy preservation is proposed. The approach aims to select appropriate QID based on re-identification risk in each attribute, reducing subjective judgment and minimizing data loss. The methodology involves data processing, frequency calculation, re-identification risk score computation, QID selection, and evaluation using data anonymization tool used for privacy-preserving data processing and analysis. The selection process considers the likelihood of information being disclosed through a single record in each attribute based on an attack where attackers are aware of the victim’s background or a randomly selected victim whose information is available in a published or external dataset. The effectiveness of the approach is evaluated on two datasets: the adult dataset and the bank marketing dataset. The results demonstrate a significant reduction in re-identification risks and success rates of attacker models after preserving QID attributes. The proposed approach offers a valuable solution for privacy preservation while minimizing data loss and maintaining data utility.