Vision Transformer Hybrid Model for Enhanced Image Generation: Integrating VQVAE and Swin
摘要
Although transformers have achieved remarkable performance across different vision tasks, they have not yet demonstrated the comparable proficiency to CNN (ConvNets) in image generation. This paper presents an innovative Swin Transformer plus VQVAE model for encoding and decoding. Elevating Swin Transformer’s scalability and global dependency modeling, our proposed model addresses these shortcomings. This advancement fills the gap in transformer-based solutions engineered specifically for image generation thereby offering improved quality and efficiency in generating images.