<p>Data-Centric AI (DCAI) aims to improve AI through better data. Data reprogramming, a key task in DCAI, enhances data representation and has achieved strong performance in predictive tasks. However, existing methods risk privacy leakage, as sensitive features may still be inferred from reprogrammed data. To address this, we propose a privacy-preserving data reprogramming framework that transforms data representations from a generative modeling perspective. Our method consists of two phases: (1) privacy-aware knowledge acquisition via an information bottleneck-guided reinforcement learning system that extracts feature sequences as a knowledge base, and (2) privacy-preserving feature space generation, where a generative model encodes the knowledge into a latent space and identifies optimal representations through projected gradient ascent, balancing predictive power and privacy. We validate our framework on eight real-world datasets, demonstrating its ability to achieve strong performance while safeguarding privacy.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Privacy-preserving data reprogramming

  • Haoyue Bai,
  • Wangyang Ying,
  • Nanxu Gong,
  • Xinyuan Wang,
  • Yanjie Fu

摘要

Data-Centric AI (DCAI) aims to improve AI through better data. Data reprogramming, a key task in DCAI, enhances data representation and has achieved strong performance in predictive tasks. However, existing methods risk privacy leakage, as sensitive features may still be inferred from reprogrammed data. To address this, we propose a privacy-preserving data reprogramming framework that transforms data representations from a generative modeling perspective. Our method consists of two phases: (1) privacy-aware knowledge acquisition via an information bottleneck-guided reinforcement learning system that extracts feature sequences as a knowledge base, and (2) privacy-preserving feature space generation, where a generative model encodes the knowledge into a latent space and identifies optimal representations through projected gradient ascent, balancing predictive power and privacy. We validate our framework on eight real-world datasets, demonstrating its ability to achieve strong performance while safeguarding privacy.