Dost: a dual optimization method for text-guided face images style transfer
摘要
Style transfer is an important research area in computer vision, and can be divided into two categories: image-guided and text-guided. Image-guided methods require the users to provide a style reference image, which takes effort to find an image that accurately describes the desired style. In contrast, text-guided methods allow users to use text to describe style requirements more flexibly, but existing methods often generate results that are less correlated to the desired style. To address these issues, we propose DOST, a novel dual optimization method for text-guided face image style transfer. Our DOST consists of two stages. In Stage 1 (the StyleGAN2 Optimization stage), we train a generation network to synthesize stylized face images according to the given natural language prompt within a few minutes of finetuning. In Stage 2 (the Mapper Optimization stage), we propose to improve stylization performance by explicitly incorporating the textural features of text prompt into the model inputs with a cross-modal multi-level mapper network. Besides, we propose a progressive strategy for training the mapper network and introduce a new patch-wise AugCLIP to constrain the style consistency between the output stylized image and the text prompt. Extensive experiments show that our DOST outperforms the state-of-the-art competitors.