FoldGEN: Multimodal Transformer for Garment Sketch-to-Photo Generation
摘要
Garment sketch-to-photo generation is one of the most crucial steps in garment design. Most existing methods contain only single conditional information, and it is challenging to combine multiple condition information. At the same time, these methods cannot generate garment folds based on sketch strokes and face a low-fidelity problem. Therefore, this paper proposes a two-stage multi-modal framework, FoldGEN, to generate garment images with folds using sketches and descriptive text as conditional information. In the first stage, we combine feature matching of discriminators and semantic perception of Convolutional Neural Network in vector quantization, which can reconstruct the details and folds of the garment images. In the second stage, a multi-conditional constrained Transformer is used to establish the association between different modality data, which allows the generated images to contain not only text description information but also folds corresponding to the strokes of the sketch. Experiments show that our method can generate garment images with different folds from sketches with high fidelity while achieving the best FID and IS on both unimodal and multi-modal tasks.