Embedding customized triggers in frequency domain for simplified and efficient detection of deep neural network backdoors
摘要
Deep neural networks (DNNs) are vulnerable to backdoor attacks, where the model produces incorrect outputs when presented with inputs containing specific triggers. However, existing trigger detection methods often suffer from high implementation complexity and significant computational overhead. This paper proposes a novel and lightweight approach by embedding customized triggers in the frequency domain to effectively detect DNN backdoors. By embedding a custom frequency-domain trigger into potentially untrustworthy images and analyzing the model’s responses, it becomes possible to differentiate between clean and poisoned inputs. Experimental results demonstrate that our method achieves average detection rates of 97.47%, 99.40%, and 99.31% on the Fashion-MNIST, CIFAR-10, and EuroSAT-RGB datasets, respectively. This efficient detection strategy provides a practical and scalable solution to enhance DNN security against backdoor attacks.