错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Identifying and Mitigating Bias in AI-Generated Image Datasets for Better Cognitive Understanding

  • Aboli Marathe,
  • Aditya Desai,
  • Rahee Walambe,
  • Ketan Kotecha

摘要

Cognitive behaviour and the effect of synthetically generated data are highly correlated. Also, human cognition is highly influenced by the bias in the data – be it visual, audio or textual data. With the new and improved tools for generating AI content, it is very challenging to ensure unbiased cognitive understanding. To that end, in this work, we focus on the image data which synthetically generated and how bias creeps into it. Modern text-to-image models with their ever-growing capabilities have been gaining traction for the quality, text-faithfulness and perceived “realism” observed in their generated outputs. Do these models truly produce realistic creations? Is an underlying bias inadvertently present in these synthetic datasets? And how can we assess the “realism” of these images? Exploring these questions and their solution is the key contribution of this paper. We propose a series of experiments for understanding the human and artificial perceptions of “realism” in synthetic data and the real world. Additionally, open questions like unprompted gender and age bias present in synthetic datasets are investigated and techniques for mitigating similar concerns in future datasets are addressed. The WEDGE dataset is used for the experimentation. Our results show that realism classification is highly dependent on human annotations within the distributions in Fig. 1. We found strong age bias in selected synthetic data and proposed simple state-of-the-art generative methods for mitigation of age, gender and realism bias.