Enhancing Virtual Try-On with Text-Image Fusion Guidance
摘要
Virtual Try-On (VTON) aims to generate authentic try-on results while preserving the identity of the person and maintaining fine-grained garment details. The core challenge of the VTON task lies in insufficient preservation of garment semantics while synthesizing garment onto the human body. To tackle this challenge, we introduce TIF-VTON, a text-image fusion guidance virtual try-on model that leverages the power of pretrained latent diffusion models. Specifically, we employ the semantic parsing module of the Qwen-VL large model to generate fine-grained textual descriptions, capturing semantic attributes of the clothing such as fabric texture and style, thereby compensating for the shortcomings of brief textual prompts in previous methods in terms of detail expression and semantic completeness, and providing more precise semantic guidance. Meanwhile, we incorporate the CLIP visual encoder for image feature extraction, introducing rich spatial and structural information from the image to further enhance the authenticity and quality of the generated results. In addition, we employ an enhanced cross-attention mechanism to jointly guide the generation of try-on images using both text and image inputs. A series of comprehensive experiments on multiple widely-used datasets show that our approach outperforms prior diffusion-based and GAN-based approaches, achieving superior performance.