The rapid evolution of artificial intelligence, propelled by large models, has led to the widespread integration of deep learning applications into everyday life. However, this convenience comes with significant security risks, especially in terms of model security. Data poisoning attacks, including targeted clean-label data poisoning attacks, pose severe threats by manipulating model outputs through covertly injected malicious poisoned samples into the training set. Despite efforts to generate more effective poisoned samples, a crucial question remains unanswered: are different target samples equally vulnerable to such attacks, and how can adversaries measure this vulnerability? To address this gap, we analyze the vulnerability of different target samples in targeted clean-label data poisoning attacks. Through experiments, we identify variations in vulnerability and demonstrate that selecting inappropriate target samples increases attack difficulty to some extent. Furthermore, we propose the Targeted Boundary Drift Velocity(TBDV), a novel measurement based on the velocity of the poisoned decision boundary over continuous time. Our experiments reveal that TBDV aids attackers in precisely identifying vulnerable samples, consequently enhancing attack success rates or diminishing the number of poisoned samples necessary in both single-target and multi-target scenarios.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Analyzing the Vulnerabilities of Targets in Clean-Label Data Poisoning Attacks

  • Yaoyu Jin,
  • Xiaochun Yang,
  • Jian Li,
  • Rong Pu,
  • Yujie Wang,
  • Bin Wang

摘要

The rapid evolution of artificial intelligence, propelled by large models, has led to the widespread integration of deep learning applications into everyday life. However, this convenience comes with significant security risks, especially in terms of model security. Data poisoning attacks, including targeted clean-label data poisoning attacks, pose severe threats by manipulating model outputs through covertly injected malicious poisoned samples into the training set. Despite efforts to generate more effective poisoned samples, a crucial question remains unanswered: are different target samples equally vulnerable to such attacks, and how can adversaries measure this vulnerability? To address this gap, we analyze the vulnerability of different target samples in targeted clean-label data poisoning attacks. Through experiments, we identify variations in vulnerability and demonstrate that selecting inappropriate target samples increases attack difficulty to some extent. Furthermore, we propose the Targeted Boundary Drift Velocity(TBDV), a novel measurement based on the velocity of the poisoned decision boundary over continuous time. Our experiments reveal that TBDV aids attackers in precisely identifying vulnerable samples, consequently enhancing attack success rates or diminishing the number of poisoned samples necessary in both single-target and multi-target scenarios.