<p>The rapid growth of AI-generated images in fields such as entertainment, e-commerce, and media has heightened the demand for robust evaluation methods to ensure high-quality, photorealistic outputs. However, current computational metrics often lack alignment with human perception, creating a gap in accurately assessing the quality of AI-generated visuals. This study introduces subjective human assessments named Visual Verity, alongside objective computational metrics, to evaluate photorealism, image quality, and text-image alignment in AI-generated images. We designed a comprehensive questionnaire and benchmarked these assessments against human judgments. The experiments are conducted using state-of-the-art models, including DALL<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="530_2025_1769_Article_IEq1.gif" Format="GIF" Height="9" Rendition="HTML" Resolution="72" Type="Linedraw" Width="8" /> </InlineMediaObject> <EquationSource Format="TEX">\(\cdot\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>·</mo> </math></EquationSource> </InlineEquation>E 2, DALL<InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="530_2025_1769_Article_IEq1.gif" Format="GIF" Height="9" Rendition="HTML" Resolution="72" Type="Linedraw" Width="8" /> </InlineMediaObject> <EquationSource Format="TEX">\(\cdot\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>·</mo> </math></EquationSource> </InlineEquation>E 3, GLIDE, and Stable Diffusion by comparing their outputs with camera-generated images. Our findings show that while AI models excel in image quality, camera-generated images surpass them in photorealism and text-image alignment. Further analysis benchmarks traditional metrics, such as SSIM and PSNR, against human judgments and highlights the Interpolative Binning Scale as a more interpretable approach to metric scores. The framework provides a structured pathway for advancing the evaluation of AI-generated images and informing future developments in AI-driven visual media.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Towards a unified evaluation framework: integrating human perception and metrics for AI-generated images

  • Memoona Aziz,
  • Umair Rehman,
  • Muhammad Umair Danish,
  • Syed Ali,
  • Amir Zaib Abbasi

摘要

The rapid growth of AI-generated images in fields such as entertainment, e-commerce, and media has heightened the demand for robust evaluation methods to ensure high-quality, photorealistic outputs. However, current computational metrics often lack alignment with human perception, creating a gap in accurately assessing the quality of AI-generated visuals. This study introduces subjective human assessments named Visual Verity, alongside objective computational metrics, to evaluate photorealism, image quality, and text-image alignment in AI-generated images. We designed a comprehensive questionnaire and benchmarked these assessments against human judgments. The experiments are conducted using state-of-the-art models, including DALL \(\cdot\) · E 2, DALL \(\cdot\) · E 3, GLIDE, and Stable Diffusion by comparing their outputs with camera-generated images. Our findings show that while AI models excel in image quality, camera-generated images surpass them in photorealism and text-image alignment. Further analysis benchmarks traditional metrics, such as SSIM and PSNR, against human judgments and highlights the Interpolative Binning Scale as a more interpretable approach to metric scores. The framework provides a structured pathway for advancing the evaluation of AI-generated images and informing future developments in AI-driven visual media.