The popularity of text-to-image generative models goes with the potential danger of intentional attacks from input texts that might lead to misleading image generation. Recently, several works have studied the robustness of these models by automatically designing adversarial prompts. Prior work, namely RIATIG, utilizes a genetic algorithm evolving from an irrelevant prompt to craft a visually natural attack prompt that generates a desired target image without explicitly describing it to avoid detection. This method avoids being semantically similar to the target prompt that describes the target image controlled by a fixed threshold. However, many attack scenarios require different ranges of stealthiness represented by distinctive threshold values of similarity. Quality-Diversity optimization which searches for diverse high-quality solutions is a natural fit for this problem. We propose applying MAP-Elites, a Quality-Diversity optimization method, to seek diverse adversarial texts that are versatile for different types of stealthiness in black-box settings. Experimental results on three widely-used generative models suggest that our method successfully finds various adversarial prompts with similarities to target and initial texts spread out in many values, allowing transfer to different attack settings.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Diverse Adversarial Samples for Text-to-Image Generation via Quality-Diversity Optimization

  • Thai Huy Nguyen,
  • Ngoc Hoang Luong

摘要

The popularity of text-to-image generative models goes with the potential danger of intentional attacks from input texts that might lead to misleading image generation. Recently, several works have studied the robustness of these models by automatically designing adversarial prompts. Prior work, namely RIATIG, utilizes a genetic algorithm evolving from an irrelevant prompt to craft a visually natural attack prompt that generates a desired target image without explicitly describing it to avoid detection. This method avoids being semantically similar to the target prompt that describes the target image controlled by a fixed threshold. However, many attack scenarios require different ranges of stealthiness represented by distinctive threshold values of similarity. Quality-Diversity optimization which searches for diverse high-quality solutions is a natural fit for this problem. We propose applying MAP-Elites, a Quality-Diversity optimization method, to seek diverse adversarial texts that are versatile for different types of stealthiness in black-box settings. Experimental results on three widely-used generative models suggest that our method successfully finds various adversarial prompts with similarities to target and initial texts spread out in many values, allowing transfer to different attack settings.