Data Governance for Artificial Intelligence
摘要
In 2021, generative artificial intelligence technology, represented by large models, swept the world, bringing revolutionary changes to human production and life. The development of artificial intelligence has shifted from “model-centric” to “data-centric”. The theory posits that good artificial intelligence requires high-quality, large-scale, and diverse data. However, in practice, data scientists often encounter issues such as data security and privacy breaches, biased and discriminatory content output, and the problem of “high volume but low quality” data. If these issues are left unregulated, they will hinder the further development of artificial intelligence technology and may even endanger the safety of individuals, enterprises, and even national security. To address these challenges and develop more responsible and controllable artificial intelligence applications, this paper clarifies the conceptual definition of data governance for artificial intelligence (DG4AI) and proposes the main stages of artificial intelligence data governance, which are divided into nine stages including data collection and preprocessing, and the governance objects are divided into seven categories such as multimodal data and annotated data. This paper also proposes to integrate the DataOps philosophy into the data governance steps, providing enterprises with a more efficient way of data management.