The need for high-quality tabular data synthesis (TDS) is growing in the government and social sectors, where balancing privacy and utility is critical. These sectors often manage large, sensitive datasets, posing challenges for effective data use without violating privacy. This survey reviews synthetic data generation methods specifically designed for public sector applications, emphasizing their strengths, weaknesses, and trade-offs. Unlike general TDS surveys, this review focuses on the unique challenges of government data, where privacy, utility, computational efficiency, and model transparency are equally vital. We identify limitations in current TDS generation and evaluation approaches and propose three key research directions to address these challenges: (1) developing unified evaluation metrics to quantify the trade-off between data utility and privacy, (2) leveraging ensemble models for more efficient and scalable synthetic data generation, and (3) creating a task-driven framework to guide the selection of appropriate TDS algorithms based on specific use cases. This review reflects the importance of tailored, ethically responsible solutions to meet the unique demands of public sector data applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Generative AI for Tabular Data Synthesis

  • Alex X. Wang,
  • Binh P. Nguyen,
  • Colin R. Simpson

摘要

The need for high-quality tabular data synthesis (TDS) is growing in the government and social sectors, where balancing privacy and utility is critical. These sectors often manage large, sensitive datasets, posing challenges for effective data use without violating privacy. This survey reviews synthetic data generation methods specifically designed for public sector applications, emphasizing their strengths, weaknesses, and trade-offs. Unlike general TDS surveys, this review focuses on the unique challenges of government data, where privacy, utility, computational efficiency, and model transparency are equally vital. We identify limitations in current TDS generation and evaluation approaches and propose three key research directions to address these challenges: (1) developing unified evaluation metrics to quantify the trade-off between data utility and privacy, (2) leveraging ensemble models for more efficient and scalable synthetic data generation, and (3) creating a task-driven framework to guide the selection of appropriate TDS algorithms based on specific use cases. This review reflects the importance of tailored, ethically responsible solutions to meet the unique demands of public sector data applications.