The process of colorizing visual media has existed since the inception of photography, as color is a crucial element that enhances the quality and appeal of such representations. The advent of deep learning techniques has led to significant advancements in video colorization, with the development of deep learning video colorization (DLVC) gaining prominence. These solutions typically entail the introduction of novel architectures or methodologies that integrate color and temporal information from videos. However, these approaches do not align with the self-supervised learning methodologies prevalent in the computer vision community, such as contrastive and autoencoder methods. In this paper, we propose a novel framework for training deep learning video colorization (DLVC) models. To the best of our knowledge, this is the first self-supervision training framework in this domain. To validate our framework, we implemented two distinct architectures, ViT and ResNet50, which were trained on the DAVIS, LDV, and UVO datasets. The results obtained from this study demonstrated superior performance on these datasets when compared with state-of-the-art models in the DLVC literature.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CVC: Contrastive Autoencoders for Video Colorization

  • Leandro Stival,
  • Ricardo da Silva Torres,
  • Helio Pedrini

摘要

The process of colorizing visual media has existed since the inception of photography, as color is a crucial element that enhances the quality and appeal of such representations. The advent of deep learning techniques has led to significant advancements in video colorization, with the development of deep learning video colorization (DLVC) gaining prominence. These solutions typically entail the introduction of novel architectures or methodologies that integrate color and temporal information from videos. However, these approaches do not align with the self-supervised learning methodologies prevalent in the computer vision community, such as contrastive and autoencoder methods. In this paper, we propose a novel framework for training deep learning video colorization (DLVC) models. To the best of our knowledge, this is the first self-supervision training framework in this domain. To validate our framework, we implemented two distinct architectures, ViT and ResNet50, which were trained on the DAVIS, LDV, and UVO datasets. The results obtained from this study demonstrated superior performance on these datasets when compared with state-of-the-art models in the DLVC literature.