错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adversarial Example Detection Based on Semantic Deviation

  • XianXian Zhang,
  • BaoPing Li

摘要

Adversarial attacks have been verified to pose a substantial threat to the artificial intelligence security. Detecting adversarial example is a primary means of achieving defense against adversarial attacks. However, the majority of methods focus on input-level vulnerabilities or distribution differences, yet they struggle to handle diverse attacks and perturbations effectively. In consideration of the abnormal semantic deviation between adversarial examples and their reconstructed images, a new method for detecting adversarial examples is proposed. Specifically, adversarial examples is detected in the model's feature space using a reconstruction mechanism combined with an improved VICReg representation learning framework and logits distribution reshaping. This approach can effectively distinguish adversarial examples when their representations violate VICReg distribution constraints or show reduced cosine similarity during reconstruction. Experiments show strong robustness against multiple attack types and perturbation levels, while improving clean sample recognition accuracy compared to existing methods.