Data Chameleon: A Self-adaptive Synthetic Data Management System
摘要
The data economy thrives on data-centric collaboration between organizations. However, open data sharing remains a pipe dream without addressing pragmatic, regulatory, and strategic concerns – which include data protection and confidentiality. Generative artificial intelligence supports the production of realistic synthetic datasets on demand, and is a promising technology to alleviate such concerns. However, the replacement of an original dataset with a synthetic dataset incurs a specific trade-off between utility and privacy, and the appropriateness of this trade-off is highly context- and application-dependent. Manually establishing and managing different synthetic generators that have diverging properties is error-prone and time-consuming, lacks flexibility, and thus is costly and impractical. This paper introduces Data Chameleon, a novel self-adaptive data management architecture for different synthetic data generators. Data Chameleon adaptively samples from different synthetic data generators in function of the data request at hand. Furthermore, in a longer-term adaptation loop, Data Chameleon monitors and evaluates the overall suitability of the available generators, to monitor evolutions in data demand, or possible concept drifts. Based on this, the Chameleon autonomously decides to re-train existing generators, or instantiate additional ones in a self-adaptive manner. The Data Chameleon architecture enhances the practical applicability of synthetic data generation, enabling more efficient and secure data sharing in real-world scenarios.