Enhancing Privacy Through Cluster-Based Anonymization of Quasi Identifiers in Correlated Datasets
摘要
In the contemporary landscape, the pervasive integration of machine learning and artificial intelligence into our daily routines has resulted in an increased focus on data-centric living. The optimal functionality of these algorithms necessitates the utilization of authentic, voluminous, diverse, and intricate datasets as their primary inputs. However, the reluctance of the general public to furnish such data stems from apprehensions regarding privacy. Thus, there arises a pressing need to devise a robust framework capable of safeguarding data privacy without compromising its utility. Thus, within the framework of our current research, we present an intelligent model meticulously engineered to tackle the intricate privacy challenges inherent in data management. Drawing upon the innate capabilities of artificial intelligence, machine learning, and data correlation analysis, our approach aims to anonymize not only sensitive attributes but also Quasi Identifiers, thereby enhancing privacy protection. Furthermore, our proposed model is intricately tailored to accommodate varying levels of data utility, ensuring that the data remains practically useful. Through a meticulously conducted series of experiments and evaluations, our hybrid model demonstrates superiority over existing algorithms, as evidenced by its outstanding performance, primarily assessed using the Information Gain metric.