Data Minimization Challenges in General Purpose AI Models
摘要
With AI becoming increasingly pervasive across various sectors, the concept of data minimization becomes crucial not just for compliance with data protection laws but as a foundational element of ethical AI development. We concede that this is not a popular opinion. Data privacy has not been among the top considerations when developing Large Language Models (LLMs). But we think it is paramount to resist as a new reality the practices that have become ubiquitous over the past few years (i.e., incessant PII scraping for generative AI model development with a disregard for the long-standing privacy principles this violates).