Stealthily Launch Backdoor Attacks Against Deep Neural Network Models via Steganography
摘要
The widespread deployment of deep neural network (DNNs) model based on image classification in the real world provides new attack scenarios for attackers. Backdoor attack is one of the most frequently used attack methods against DNNs models due to their simplicity and effectiveness. In this paper, we attempt to launch a backdoor attack by Hide Text Trigger via Steganography (HTTS). Specifically, we propose a regular text trigger and use backdoor steganography we designed to embed the trigger in a small number of training images. When launching a backdoor attack, our triggers will be mistaken by DNNs models as hidden features of the training images. DNNs models can easily capture and learn these hidden features. In the testing phase, only the same trigger embedded in the training image can activate the backdoor of the DNNs model, causing the DNNs model to output incorrect classification results. Our experiments on the CIFAR-10 data set show that https can not only achieve an attack success rate of up to 100% but also effectively alleviate the contradiction between visual invisibility and attack performance.