<p>Existing fusion-based few-shot image generation methods struggle to synthesize high-quality detailed texture features during the image generation process, resulting in a degradation in the quality of the generated images. To solve this problem, we propose a Discrete Cosine Transform Generative Adversarial Network (DCTGAN), which is based on the concept of frequency decomposition. We utilize discrete cosine transform to decompose image features into high-frequency and low-frequency components. Then, we utilize skip connections and inverse transformation to enhance the texture features and overall structure of the generated image, resulting in a substantial improvement in image quality and detail restoration capability. Additionally, we employ frequency MSE loss to facilitate the generation network in learning the subtle changes and details of frequency domain features, hence generating more realistic and clearer images. Comprehensive experiments on three general datasets demonstrate that DCTGAN exhibits significant advantages and effects in image generation tasks, successfully enhancing the quality and fidelity of generated images. New FID and LPIPS evaluation results are acquired from the Flower, Animal Faces, and VGGFace datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DCTGAN: frequency decomposition GAN for few-shot image generation

  • Jinhang Zhang,
  • Min Gao,
  • Dan Fang,
  • Chaowang Li,
  • Hongyun Wang

摘要

Existing fusion-based few-shot image generation methods struggle to synthesize high-quality detailed texture features during the image generation process, resulting in a degradation in the quality of the generated images. To solve this problem, we propose a Discrete Cosine Transform Generative Adversarial Network (DCTGAN), which is based on the concept of frequency decomposition. We utilize discrete cosine transform to decompose image features into high-frequency and low-frequency components. Then, we utilize skip connections and inverse transformation to enhance the texture features and overall structure of the generated image, resulting in a substantial improvement in image quality and detail restoration capability. Additionally, we employ frequency MSE loss to facilitate the generation network in learning the subtle changes and details of frequency domain features, hence generating more realistic and clearer images. Comprehensive experiments on three general datasets demonstrate that DCTGAN exhibits significant advantages and effects in image generation tasks, successfully enhancing the quality and fidelity of generated images. New FID and LPIPS evaluation results are acquired from the Flower, Animal Faces, and VGGFace datasets.