We propose a data-driven approach for context-aware person image generation. Specifically, we attempt to generate a novel person image such that the synthesized instance can blend into a complex scene. In our method, the position, scale, and appearance of the generated person instance are semantically conditioned on the existing persons in the scene. The proposed technique consists of three sequential steps. At first, an image-to-image translation model infers a coarse semantic mask that represents the new person’s spatial location, scale, and potential pose. Next, we introduce a data-centric approach to select the closest representation from a precomputed cluster of fine semantic masks. Finally, we use a multi-scale, attention-guided rendering network to transfer the appearance attributes from an exemplar image. The proposed strategy enables us to synthesize high-quality, semantically coherent, realistic human instances that can blend into an existing scene without altering the global context. We conclude our findings with relevant qualitative and quantitative evaluations.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Semantically Consistent Person Image Generation

  • Prasun Roy,
  • Saumik Bhattacharya,
  • Subhankar Ghosh,
  • Umapada Pal,
  • Michael Blumenstein

摘要

We propose a data-driven approach for context-aware person image generation. Specifically, we attempt to generate a novel person image such that the synthesized instance can blend into a complex scene. In our method, the position, scale, and appearance of the generated person instance are semantically conditioned on the existing persons in the scene. The proposed technique consists of three sequential steps. At first, an image-to-image translation model infers a coarse semantic mask that represents the new person’s spatial location, scale, and potential pose. Next, we introduce a data-centric approach to select the closest representation from a precomputed cluster of fine semantic masks. Finally, we use a multi-scale, attention-guided rendering network to transfer the appearance attributes from an exemplar image. The proposed strategy enables us to synthesize high-quality, semantically coherent, realistic human instances that can blend into an existing scene without altering the global context. We conclude our findings with relevant qualitative and quantitative evaluations.