DCT-SwinGAN: Leveraging DCT and Swin Transformer for Face Synthesis from Sketch and Thermal Domains
摘要
Face generation remains a crucial task owing to its applications in crime investigation, entertainment, etc. Sketch to face synthesis is an important task accomplished using Generative Adversarial Network (GAN) models. The generator of the GAN is usually designed using Convolutional Neural Networks (CNNs). GANs are able to deliver promising outcomes in some cases but fail in others. Owing to the local nature of CNN, it is unable to keep track of long range dependencies, limiting the model to local information and delivering underwhelming results. However, recently developed Transformers encode the global context having long range dependencies. Swin Transformer is designed to be suitable for images with promising performance on several vision tasks. We utilize the Swin Transformer along with Discrete Cosine Transform (DCT) in the GAN framework and propose DCT-SwinGAN for face synthesis from sketch and thermal domains. DCT-SwinGAN comprises a multi scale discriminator paired with a generator comprising an attention module, DCT ResNet Convolution, Deconv decoder, and Swin Transformer. The proposed model captures not only local but also global information to generate realistic face images. The generated model outperforms the existing state-of-the-art models on CUHK sketch-to-face and the WHU-IIP thermal-to-face datasets.