<p>In this work we develop CAST-GAN, a Compression Aware Spatio Temporal Transformer GAN for the perceptual restoration of VVC encoded compressed video. Our approach utilizes the QP and MV information from the VVC bit stream in conjunction with a spatio-temporal transformer network based on the Swin Transformer architecture, to generate restored video frames. A compression-aware multi-branch discriminator is introduced to enforce spatial fidelity, motion consistency, and artifact suppression. The loss function utilized by our framework includes a combination of pixel-wise, adversarial, and perceptually motivated terms. In addition to these, we include a term that encourages temporal coherence in order to address the issues related to flicker and temporal inconsistency in restored videos. Experiments conducted on Vimeo-90&#xa0;K, REDS, UVG, and BVI-DVC datasets demonstrate that CAST-GAN outperforms SwinIR, VRT, and BasicVSR + + across distortion, perceptual, and temporal metrics. The CAST-GAN demonstrates a significant improvement in PSNR (+ 0.8–1.3 dB), a 38% decrease in LPIPS, and significant improvements in NIQE. Our results are obtained while maintaining real time processing rates (45 fps at 720p) measured using batch size 1 during inference. The results of ablation studies provide evidence supporting the utility of QP conditioning and compression aware adversarial training in achieving better quality restored video. These results suggest the potential for CAST-GAN to be used in real-time video enhancement applications as well as adaptive streaming.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CAST-GAN: compression-aware spatio-temporal transformer GAN for perceptual enhancement of VVC-compressed video

  • L. Balaji,
  • A. Dhanalakshmi,
  • B. V. Santhosh Krishna,
  • Roshan Fernandes,
  • G. Elumalai

摘要

In this work we develop CAST-GAN, a Compression Aware Spatio Temporal Transformer GAN for the perceptual restoration of VVC encoded compressed video. Our approach utilizes the QP and MV information from the VVC bit stream in conjunction with a spatio-temporal transformer network based on the Swin Transformer architecture, to generate restored video frames. A compression-aware multi-branch discriminator is introduced to enforce spatial fidelity, motion consistency, and artifact suppression. The loss function utilized by our framework includes a combination of pixel-wise, adversarial, and perceptually motivated terms. In addition to these, we include a term that encourages temporal coherence in order to address the issues related to flicker and temporal inconsistency in restored videos. Experiments conducted on Vimeo-90 K, REDS, UVG, and BVI-DVC datasets demonstrate that CAST-GAN outperforms SwinIR, VRT, and BasicVSR + + across distortion, perceptual, and temporal metrics. The CAST-GAN demonstrates a significant improvement in PSNR (+ 0.8–1.3 dB), a 38% decrease in LPIPS, and significant improvements in NIQE. Our results are obtained while maintaining real time processing rates (45 fps at 720p) measured using batch size 1 during inference. The results of ablation studies provide evidence supporting the utility of QP conditioning and compression aware adversarial training in achieving better quality restored video. These results suggest the potential for CAST-GAN to be used in real-time video enhancement applications as well as adaptive streaming.