Backdoor attacks typically attempt to import poisoned samples during model training to mislead the model into making incorrect predictions. The concealment of backdoor attacks makes it difficult to defend against them, as the attacked model only becomes abnormal when encountering poisoned samples. Defenders typically need knowledge related to attack pattern in poisoned samples to adjust the model to prevent backdoor attacks. The existing backdoor defense works make adversarial assumption about the backdoor attack pattern and optimize complex adversarial perturbation as the backdoor attack pattern. However, the complex adversarial perturbation optimization process will introduce various optimization objectives and increase computational consumption. To address above issue, we propose a fine-tuning backdoor defense method based on scaled prediction consistency intermediate-level adversarial sample selection called FT-SPC. In FT-SPC, several candidate samples are obtained through intermediate-layer attacks and the scaled prediction consistency is used to search for samples that generalize well to poisoned samples. The experiments show that our approach is comparable to existing fine-tuning defense methods in terms of defense effectiveness rate. Ablation studies demonstrate the effectiveness of the selection strategy based on scaled prediction consistency in enhancing model defense.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FT-SPC: A Fine-Tuning Approach for Backdoor Defense via Adversarial Sample Selection

  • Shichang Chen,
  • Ya Li

摘要

Backdoor attacks typically attempt to import poisoned samples during model training to mislead the model into making incorrect predictions. The concealment of backdoor attacks makes it difficult to defend against them, as the attacked model only becomes abnormal when encountering poisoned samples. Defenders typically need knowledge related to attack pattern in poisoned samples to adjust the model to prevent backdoor attacks. The existing backdoor defense works make adversarial assumption about the backdoor attack pattern and optimize complex adversarial perturbation as the backdoor attack pattern. However, the complex adversarial perturbation optimization process will introduce various optimization objectives and increase computational consumption. To address above issue, we propose a fine-tuning backdoor defense method based on scaled prediction consistency intermediate-level adversarial sample selection called FT-SPC. In FT-SPC, several candidate samples are obtained through intermediate-layer attacks and the scaled prediction consistency is used to search for samples that generalize well to poisoned samples. The experiments show that our approach is comparable to existing fine-tuning defense methods in terms of defense effectiveness rate. Ablation studies demonstrate the effectiveness of the selection strategy based on scaled prediction consistency in enhancing model defense.