Model stealing allows to extract the knowledge of a deployed target machine learning model by sending query images and training a substitute model via the inference results. However, accessing the same data used to train the target model is often impractical. Recent data-free model stealing methods overcome this limit using specifically crafted noise as queries. However, two major flaws limit their effectiveness. First, data-free methods suffer from low query efficiency, requiring high query budgets, such as over 20 million queries to train a substitute for CIFAR-10. Second, crafted queries lack any perceptual semantics. Hence, they can easily be filtered out from legitimate requests. To address these issues, we propose AEDM, a framework for Adversarial knowledge Extraction via steering Diffusion Models. AEDM leverages publicly available pre-trained diffusion models to craft adversarial query images, controlled through the latent variable fed into the diffusion models, to maximize the knowledge transferred to the substitute model. These queries not only resemble in all aspects legitimate requests of images with a semantic meaning, but also significantly lower the number of required queries. Results on three datasets demonstrate that, given a budget as limited as 1/200 queries of the baselines, the accuracy of our trained substitute model can outperform that of state-of-the-art data-free stealing methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adversarial Knowledge Extraction via Steering Diffusion Models

  • Chi Hong,
  • Jiyue Huang,
  • Lydia Chen,
  • Robert Birke

摘要

Model stealing allows to extract the knowledge of a deployed target machine learning model by sending query images and training a substitute model via the inference results. However, accessing the same data used to train the target model is often impractical. Recent data-free model stealing methods overcome this limit using specifically crafted noise as queries. However, two major flaws limit their effectiveness. First, data-free methods suffer from low query efficiency, requiring high query budgets, such as over 20 million queries to train a substitute for CIFAR-10. Second, crafted queries lack any perceptual semantics. Hence, they can easily be filtered out from legitimate requests. To address these issues, we propose AEDM, a framework for Adversarial knowledge Extraction via steering Diffusion Models. AEDM leverages publicly available pre-trained diffusion models to craft adversarial query images, controlled through the latent variable fed into the diffusion models, to maximize the knowledge transferred to the substitute model. These queries not only resemble in all aspects legitimate requests of images with a semantic meaning, but also significantly lower the number of required queries. Results on three datasets demonstrate that, given a budget as limited as 1/200 queries of the baselines, the accuracy of our trained substitute model can outperform that of state-of-the-art data-free stealing methods.