错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Neuron-Level Inverse Perturbation Against Adversarial Attacks

  • Jinyin Chen,
  • Ximin Zhang,
  • Haibin Zheng

摘要

Although deep learning models have achieved unprecedented success, their vulnerabilities towards adversarial attacks have attracted increasing attention, especially when deployed in security-critical domains. Numerous defense methods, including reactive and proactive ones, have been proposed for model robustness improvement. The former ones, such as conducting transformations to remove perturbations, usually fail to handle large perturbations via adaptivity. The proactive defenses that involve retraining, suffer from the issue of the attack dependency and high computation cost. Addressing these challenges, we consider defense methods from a novel perspective of the general effect of adversarial attacks that take on model neurons. Specifically, we introduce the concept of neuron influence, which is a measurement of neurons’ contribution to correct classification. Then, we observed that almost all attacks fool the model by suppressing neurons with large influence and enhancing those with small influence. Based on this pattern, we propose an attack-agnostic defense, Neuron-level inverse perturbation, achieving defense against general attacks by in turn strengthening neurons with larger influence and weakening those with smaller influence. Only a few batches of benign examples are required for neuron influence calculation and our method can cope with different perturbation sizes adaptively. Comprehensive experiments conducted on three image datasets and six models demonstrate that our method shows better defense success rate ( \(\sim \!\!\times 1.3\) ) than the state-of-the-art baselines against eleven adversarial attacks, with only 1/13 time cost.