Virtual try-on systems have gained immense popularity, offering convenience and personalized experiences for online shopping. However, many existing systems lack robust handling of diverse human poses and background settings. This study proposes a comprehensive Framework for Virtual Try-On with Background Processing and Human Pose Estimation, leveraging both pre-trained and custom-trained models. The framework commences by utilizing U2Net (Qin, Zhang, Huang, Dehghan, Zaiane, Jagersand, Pattern Recogn, 106:107404, 2020) for background segmentation, preserving the original backdrop for later use. For the human subject, the system employs pre-trained OpenPose and DensePose models for 2D and 3D pose estimation, respectively. Further semantic segmentation is accomplished using a custom-trained Mask2Former model, identifying distinct body and clothing areas. The clothing item also undergoes similar segmentation. All these processed inputs are then fed into a pre-trained GP-VTON model for the final virtual try-on. The synthesized output is seamlessly integrated with the original background, providing a realistic virtual try-on experience. Experimental results confirm the effectiveness of our framework in handling a wide range of poses and backgrounds.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Framework for Virtual Try-on with Background Processing and Human Pose Estimation

  • Chih-Chiang Lin,
  • Jhing-Fa Wang

摘要

Virtual try-on systems have gained immense popularity, offering convenience and personalized experiences for online shopping. However, many existing systems lack robust handling of diverse human poses and background settings. This study proposes a comprehensive Framework for Virtual Try-On with Background Processing and Human Pose Estimation, leveraging both pre-trained and custom-trained models. The framework commences by utilizing U2Net (Qin, Zhang, Huang, Dehghan, Zaiane, Jagersand, Pattern Recogn, 106:107404, 2020) for background segmentation, preserving the original backdrop for later use. For the human subject, the system employs pre-trained OpenPose and DensePose models for 2D and 3D pose estimation, respectively. Further semantic segmentation is accomplished using a custom-trained Mask2Former model, identifying distinct body and clothing areas. The clothing item also undergoes similar segmentation. All these processed inputs are then fed into a pre-trained GP-VTON model for the final virtual try-on. The synthesized output is seamlessly integrated with the original background, providing a realistic virtual try-on experience. Experimental results confirm the effectiveness of our framework in handling a wide range of poses and backgrounds.