Pose Guided Person Image Generation (PGPIG) aims to transform persons in source images into given target poses. Most existing methods only distort texture information towards the target pose, ignoring the impact of pose information transformation, resulting in images with unrealistic poses or texture loss. In this paper, we propose a novel generation network that sequentially conducts pose and texture transformations to enhance PGPIG performance. Initially, we prioritize pose transformation, generating features aligning with the target pose while retaining source texture details. Subsequently, we concentrate on texture transformation, ensuring consistency with the target pose. To achieve the goal of each step, we propose the Pose Factor Transformer Block (PT) for injecting target pose information and the Texture Factor Transformer Block (TT) for refining texture details in the generated person image. Extensive experiments demonstrate the efficacy of our approach across evaluation metrics such as LPIPS, PSNR, and SSIM. Furthermore, our network does not require additional parsing labels and reduces training costs significantly.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Sequential Transfer of Pose and Texture for Pose Guided Person Image Generation

  • Zifan Li,
  • Qingxuan Shi,
  • Shuishui Cheng

摘要

Pose Guided Person Image Generation (PGPIG) aims to transform persons in source images into given target poses. Most existing methods only distort texture information towards the target pose, ignoring the impact of pose information transformation, resulting in images with unrealistic poses or texture loss. In this paper, we propose a novel generation network that sequentially conducts pose and texture transformations to enhance PGPIG performance. Initially, we prioritize pose transformation, generating features aligning with the target pose while retaining source texture details. Subsequently, we concentrate on texture transformation, ensuring consistency with the target pose. To achieve the goal of each step, we propose the Pose Factor Transformer Block (PT) for injecting target pose information and the Texture Factor Transformer Block (TT) for refining texture details in the generated person image. Extensive experiments demonstrate the efficacy of our approach across evaluation metrics such as LPIPS, PSNR, and SSIM. Furthermore, our network does not require additional parsing labels and reduces training costs significantly.