错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TrojanInterpret: A Detecting Backdoors Method in DNN Based on Neural Network Interpretation Methods

  • Oleg Pilipenko,
  • Bulat Nutfullin,
  • Vasily Kostyumov

摘要

Neural networks are increasingly used in various applications, but their training most often requires a huge amount of data. Inspecting the entire training dataset becomes impossible in such a situation and creates the opportunity for backdoor attacks when an attacker injects special triggers into training data. It makes the ML model perform incorrectly in the presence of the trigger while behaving normally with clean inputs. In this paper, we present a new method for backdoor detection in neural networks. This approach is based on neural network interpretation techniques and uses the idea of the difference between distributions of saliency values of neural networks with backdoors and without them. The proposed method demonstrated robustness when detecting backdoors on several datasets.