Existing detection methods for adversarial attacks in medical images mostly rely on prior knowledge about the attacks and the target models. This work introduces a new attack detector termed as AdaSVaT, that is both image and model agnostic, employing specific noise reduction technique through Adaptive Singular Value Thresholding (ASVT). The method exploits the significant impact of adversarial attacks on the lower singular values of an image. The AdaSVaT algorithm adaptively thresholds the singular values of the input image with the help of a linear regressor to generate a low-rank version of the same. Both the original and low-rank versions are experimented with the state-of-the-art classifiers. Adversarial examples are detected by examining the classification inconsistency between the input image and its low-rank version. Additionally, experimental results on three fundus image datasets (Kaggle EyePACS, IDRID and APTOS) prove the negligible loss of information from the images during reconstruction. The proposed method achieves improved adversarial detection accuracy with minimal computational burden while maintaining high structural similarity in both binary and multi-class classification tasks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AdaSVaT: Adaptive Singular Value Thresholding for Adversarial Detection in Fundus Images

  • Nirmal Joseph,
  • Sudhish N. George,
  • P. M. Ameer,
  • Kiran Raja

摘要

Existing detection methods for adversarial attacks in medical images mostly rely on prior knowledge about the attacks and the target models. This work introduces a new attack detector termed as AdaSVaT, that is both image and model agnostic, employing specific noise reduction technique through Adaptive Singular Value Thresholding (ASVT). The method exploits the significant impact of adversarial attacks on the lower singular values of an image. The AdaSVaT algorithm adaptively thresholds the singular values of the input image with the help of a linear regressor to generate a low-rank version of the same. Both the original and low-rank versions are experimented with the state-of-the-art classifiers. Adversarial examples are detected by examining the classification inconsistency between the input image and its low-rank version. Additionally, experimental results on three fundus image datasets (Kaggle EyePACS, IDRID and APTOS) prove the negligible loss of information from the images during reconstruction. The proposed method achieves improved adversarial detection accuracy with minimal computational burden while maintaining high structural similarity in both binary and multi-class classification tasks.