Adversarial attacks manipulate Deep Neural Networks (DNNs) by introducing imperceptible noise to input samples, resulting in inaccurate predictions and posing a notable security threat to deep learning systems. Detecting adversarial samples is crucial for enhancing the robustness and security of DNNs. However, current adversarial defense detection methods rely on detection features chosen through statistical experience, lacking explainability and leading to suboptimal detection performance. To address this, our study introduces an explainable method for detecting adversarial samples. We propose an Adversarial Detection method based on Prediction Feature, named ADPF. Adversarial samples can induce misclassifications in DNNs, there are distinct disparities in the prediction features of normal versus adversarial samples. ADPF employs an attention mechanism and a meticulously designed feature extraction network for detecting inconsistent prediction features, effectively discerning between normal and adversarial samples. For detector training, ADPF utilizes one-class classifiers to train detection features, significantly enhancing the accuracy in detecting various types of adversary samples. Extensive experiments demonstrate that ADPF significantly outperforms conventional detection frameworks, achieving enhanced accuracy in identifying adversarial samples, particularly in scenarios involving feature-based adversarial attacks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ADPF: Adversarial Sample Detection Based on Prediction Feature

  • Pengju Wang,
  • Jing Liu

摘要

Adversarial attacks manipulate Deep Neural Networks (DNNs) by introducing imperceptible noise to input samples, resulting in inaccurate predictions and posing a notable security threat to deep learning systems. Detecting adversarial samples is crucial for enhancing the robustness and security of DNNs. However, current adversarial defense detection methods rely on detection features chosen through statistical experience, lacking explainability and leading to suboptimal detection performance. To address this, our study introduces an explainable method for detecting adversarial samples. We propose an Adversarial Detection method based on Prediction Feature, named ADPF. Adversarial samples can induce misclassifications in DNNs, there are distinct disparities in the prediction features of normal versus adversarial samples. ADPF employs an attention mechanism and a meticulously designed feature extraction network for detecting inconsistent prediction features, effectively discerning between normal and adversarial samples. For detector training, ADPF utilizes one-class classifiers to train detection features, significantly enhancing the accuracy in detecting various types of adversary samples. Extensive experiments demonstrate that ADPF significantly outperforms conventional detection frameworks, achieving enhanced accuracy in identifying adversarial samples, particularly in scenarios involving feature-based adversarial attacks.