Text-to-image generative models have revolutionized the creation of contextually accurate and visually realistic images. This paper evaluates the performance of prominent open-source models, including Stable Diffusion XL (SDXL), Stable Diffusion v2, and Stable Diffusion v1, for generating interior design images. The study utilizes the ADE MIT Dataset and the Hugging Face Interior Dataset to explore the preparation and fine-tuning of these models, comparing pre-trained and fine-tuned versions using techniques such as LoRA fine-tuning and Exponential Moving Average. Results demonstrate significant improvements through fine-tuning, with SDXL achieving the best performance, reducing FID scores from 3.296 to 2.579 on the ADE MIT dataset and from 5.472 to 4.429 on the Interior dataset. Although Amused-512 exhibited higher initial FID scores (11.009 and 8.676), it achieved high CLIP scores (95.133 and 95.437) and showed consistent improvement after fine-tuning. The results underline the importance of using domain-specific datasets and fine-tuning to enhance text-image alignment and overall model performance for interior design image generation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Benchmarking Generative Models for Interior Image Synthesis and Creativity

  • Fatima Saad,
  • Ibrahim Basit,
  • Imama Amjad,
  • Faezeh Soleimani

摘要

Text-to-image generative models have revolutionized the creation of contextually accurate and visually realistic images. This paper evaluates the performance of prominent open-source models, including Stable Diffusion XL (SDXL), Stable Diffusion v2, and Stable Diffusion v1, for generating interior design images. The study utilizes the ADE MIT Dataset and the Hugging Face Interior Dataset to explore the preparation and fine-tuning of these models, comparing pre-trained and fine-tuned versions using techniques such as LoRA fine-tuning and Exponential Moving Average. Results demonstrate significant improvements through fine-tuning, with SDXL achieving the best performance, reducing FID scores from 3.296 to 2.579 on the ADE MIT dataset and from 5.472 to 4.429 on the Interior dataset. Although Amused-512 exhibited higher initial FID scores (11.009 and 8.676), it achieved high CLIP scores (95.133 and 95.437) and showed consistent improvement after fine-tuning. The results underline the importance of using domain-specific datasets and fine-tuning to enhance text-image alignment and overall model performance for interior design image generation.