ADMMOA: Attribute-Driven Multimodal Optimization for Face Recognition Adversarial Attacks
摘要
Deep Neural Networks, especially face recognition (FR) models, have been shown to be vulnerable to digital and physical adversarial samples, which involve adding subtle perturbations to benign face images to deceive the models. This vulnerability poses a significant threat to the security of FR models and the collective well-being of society. To enhance the robustness of FR models against attacks, this paper aims to improve the transferability of adversarial face examples. We propose a novel approach, Attribute-Driven Multimodal Optimization Attack (ADMMOA), which leverages the conditional latent diffusion model to create adversarial images with high transferability and image quality in the latent space. Specifically, we introduce a multimodal conditional diffusion generation module that uses an adaptive significant attribute text and a dynamic semantic mask image to generate realistic images with semantic guidance of the significant attribute in the powerful inpainting process. Moreover, with the idea of gradient attack, the CLIP-augmented adaptive semantic adversarial perturbation module is introduced to further ensure the stealthiness and attack effectiveness of the generated adversarial face images. Extensive quantitative and qualitative experiments on the publicly available CelebA-HQ dataset demonstrate the superior performance of ADMMOA in improving the black-box transferability compared to the state-of-the-art methods. Particularly, our proposed ADMMOA achieves attack success rates (ASRs) of 62.40, 90.70, 43.50, and 83.30 on IR152, IRSE50, FaceNet, and MobileFace, respectively, surpassing Adv-Diffusion by 8.3%, 5.2%, 12.1%, and 5.0%.