Transferable Adversarial Attacks via Diffusion-Based Keyword Embedding and Latent Optimization
摘要
Adversarial examples expose significant security vulnerabilities in deep neural networks (DNNs), especially through cross-model transfer attacks, where examples generated for one model effectively attack others. While \({L}_{p}\) -norm-based methods have shown success in attack rates, they often generate high-frequency noise that is perceptible to the human eye. Recently, diffusion-based unrestricted methods have attracted attention for their advancements in both imperceptibility and transferability. However, existing diffusion-based methods primarily introduce perturbations in the latent space, overlooking the potential role of text embeddings in the denoising process and the enhancement of adversarial effectiveness. Furthermore, these methods lack exploration of the interaction between text guidance and latent optimization. We propose a novel framework, Diffusion-based Keyword Embedding and Latent Optimization (D-KLO) to address these limitations. After mapping images into the latent space, we optimize keyword embeddings and latents separately. By combining a composite loss function and a multi-stage adaptive restart strategy, the proposed framework generates adversarial examples with enhanced transferability. Experimental results demonstrate that D-KLO outperforms state-of-the-art attack methods, achieving 3.3%–28.6% and 6.2%–29.7% performance gains on normally trained models and defense mechanisms, respectively. This validates the effectiveness and superiority of the proposed approach.