Easy for Us, Complex for AI: Assessing the Coherence of Generated Realistic Images
摘要
The existence of several effective AI methods that generate realistic images has garnered significant attention in recent years. Notable generative approaches utilizing different AI techniques, including Variational Autoencoders (VAEs), Generative Adversarial Networks (GANs), style-based generators, transformers, and diffusion models, have become popular for producing high-quality and diverse images. However, while humans can naturally assess the realistic qualities of images, this task remains challenging for quantitative automatic systems. This document presents a survey of current metrics and criteria for assessing the coherence of AI-generated images in terms of realism and quality. This review highlights how the metrics for realism in generated images fall short of understanding the high coherence that is essential for creating a realistic and immersive visual experience indistinguishable from real-world scenes. This follows Moravec’s paradox, which states that tasks easy for humans, such as pattern recognition, are often difficult for computers, which proves that the search for realism metrics keeps going. Specifically, this review discusses how the coherence of generated realistic images contributes to their believability by human perception, showing various levels of structural consistency, harmony, and logical connections between elements such as texture, color, lighting, and temporal aspects to create a cohesive scene that aligns with our understanding of the real world.