Enforcing Right to Be Forgotten in Cloud-Based Data Lakes
摘要
This paper focuses on using metadata to enforce the right to be forgotten in large-scale data lakes. With the rise of cloud storage services for massive data storage, ensuring compliance with data protection regulations like GDPR has become challenging. Implementing the right to be forgotten in cloud-based data lakes is complex due to different storage systems and immutability properties. Existing solutions lack user specific information, emphasizing the need for a practical approach. This paper presents a novel solution that leverages metadata to address these challenges. Our solution is faster, supports various file types, and generates user-specific PII reports for deletion or anonymization. Evaluation against existing tools demonstrates its effectiveness. By leveraging metadata, our solution ensures compliance with data protection laws and overcomes the challenges of diverse storage systems in cloud-based data lakes.