A Transformer-Based Video Colorization Method Fusing Local Self-attention and Bidirectional Optical Flow
摘要
Video colorization encounters two principal challenges: colorization quality and temporal flicker. Balancing colorization quality and temporal consistency is a significant challenge. To address the aforementioned issues, we employ the transformer model in the field of video colorization and introduce a pioneering video colorization method named VCTR. Initially, to guarantee the quality of video coloring, we employ a pretrained image coloring network to add color to the grayscale video frame. Next, the feature extraction module for VCTR is utilized to perform feature extraction and propagation on the color frames during the colorization process. Finally, the transformer module is designed to fully leverage local information via the local feature self-attention layer. Additionally, motion information from bidirectional optical flow is utilized to identify correlations across video frames for feature fusion, guaranteeing both the coloring effect and temporal consistency. The experimental results demonstrate that VCTR outperforms existing methods on two publicly available datasets, namely, DAVIS and Videvo. VCTR attains the top ranking in the long-series dataset, Videvo, based on the CTBI (Colorization and Temporal-Consistency Balance Index). This achievement underscores VCTR’s ability to strike a commendable balance between colorization quality and temporal consistency.