Virtual Try-On (VTON) aims to generate authentic try-on results while preserving the identity of the person and maintaining fine-grained garment details. The core challenge of the VTON task lies in insufficient preservation of garment semantics while synthesizing garment onto the human body. To tackle this challenge, we introduce TIF-VTON, a text-image fusion guidance virtual try-on model that leverages the power of pretrained latent diffusion models. Specifically, we employ the semantic parsing module of the Qwen-VL large model to generate fine-grained textual descriptions, capturing semantic attributes of the clothing such as fabric texture and style, thereby compensating for the shortcomings of brief textual prompts in previous methods in terms of detail expression and semantic completeness, and providing more precise semantic guidance. Meanwhile, we incorporate the CLIP visual encoder for image feature extraction, introducing rich spatial and structural information from the image to further enhance the authenticity and quality of the generated results. In addition, we employ an enhanced cross-attention mechanism to jointly guide the generation of try-on images using both text and image inputs. A series of comprehensive experiments on multiple widely-used datasets show that our approach outperforms prior diffusion-based and GAN-based approaches, achieving superior performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Virtual Try-On with Text-Image Fusion Guidance

  • Jingyi Guo,
  • Pengfei Duan,
  • Chenghu Du,
  • Shengwu Xiong

摘要

Virtual Try-On (VTON) aims to generate authentic try-on results while preserving the identity of the person and maintaining fine-grained garment details. The core challenge of the VTON task lies in insufficient preservation of garment semantics while synthesizing garment onto the human body. To tackle this challenge, we introduce TIF-VTON, a text-image fusion guidance virtual try-on model that leverages the power of pretrained latent diffusion models. Specifically, we employ the semantic parsing module of the Qwen-VL large model to generate fine-grained textual descriptions, capturing semantic attributes of the clothing such as fabric texture and style, thereby compensating for the shortcomings of brief textual prompts in previous methods in terms of detail expression and semantic completeness, and providing more precise semantic guidance. Meanwhile, we incorporate the CLIP visual encoder for image feature extraction, introducing rich spatial and structural information from the image to further enhance the authenticity and quality of the generated results. In addition, we employ an enhanced cross-attention mechanism to jointly guide the generation of try-on images using both text and image inputs. A series of comprehensive experiments on multiple widely-used datasets show that our approach outperforms prior diffusion-based and GAN-based approaches, achieving superior performance.