Framework for Virtual Try-on with Background Processing and Human Pose Estimation
摘要
Virtual try-on systems have gained immense popularity, offering convenience and personalized experiences for online shopping. However, many existing systems lack robust handling of diverse human poses and background settings. This study proposes a comprehensive Framework for Virtual Try-On with Background Processing and Human Pose Estimation, leveraging both pre-trained and custom-trained models. The framework commences by utilizing U2Net (Qin, Zhang, Huang, Dehghan, Zaiane, Jagersand, Pattern Recogn, 106:107404, 2020) for background segmentation, preserving the original backdrop for later use. For the human subject, the system employs pre-trained OpenPose and DensePose models for 2D and 3D pose estimation, respectively. Further semantic segmentation is accomplished using a custom-trained Mask2Former model, identifying distinct body and clothing areas. The clothing item also undergoes similar segmentation. All these processed inputs are then fed into a pre-trained GP-VTON model for the final virtual try-on. The synthesized output is seamlessly integrated with the original background, providing a realistic virtual try-on experience. Experimental results confirm the effectiveness of our framework in handling a wide range of poses and backgrounds.