Invisible Backdoor Attacks in Image Classification Based Network Services
摘要
The chapter provides an overview of two emerging categories of stealthy backdoor attacks, their distinctions, and the two metrics of stealth level. Subsequently, two types of novel invisible backdoor attacks are introduced, including their underlying principles, detailed procedures, and experimental results. In the first type of attack, image steganography is employed to embed the trigger into the bit space of an image, thereby achieving stealthiness against system administrators while maintaining a high success rate of the backdoor attack. For the second type of attack, a regularization term is included in the optimization process to constrain the trigger generation while ensuring backdoor effectiveness, leading to generated triggers that are invisible to the human eye. In the experiment, the performance and invisibility of the two new backdoor attacks are quantitatively measured. Finally, the effectiveness of the introduced backdoor attacks against mainstream detection techniques is demonstrated.