Synthetic Versus Authentic Data
摘要
Synthetic and real-world data are used in AI training, creating potential problems. A governance framework is proposed for the ethical and successful merging of synthetic and authentic (measured) data in AI training. The framework covers technological, legal, and ethical issues to set norms and standards for synthetic data production, validation, and integration with measured data. Technically, the framework requires transparent disclosure of synthetic data proportion and characteristics, rigorous quality assurance mechanisms to validate synthetic data against real-world counterparts, and ongoing monitoring of AI models trained on fused data to detect and address biases or inconsistencies. It supports unambiguous synthetic data restrictions to comply with privacy and data protection legislation. Ethically, it emphasizes the need to overcome synthetic data biases and ensure that AI models trained on fused data do not perpetuate discrimination. By applying this governance framework, stakeholders may ensure the ethical integration of synthetic and measured data in AI training, increasing confidence in AI systems and reducing threats to individuals and society.