Multi-frame super-resolution is a process of combining multiple low-resolution observations of the same scene to reconstruct a high-resolution image. State-of-the-art techniques leverage deep networks that are trained in a supervised way to increase the similarity between the reconstruction outcome and a high-resolution reference image. The similarity is typically quantified using L1 or L2 norm when computing the loss function during training. However, these pixel-wise metrics often fail to provide adequate guidance, especially for real-world training datasets with low- and high-resolution images captured using different sensors. While exploiting perceptual similarity underpinned with deep features appears to be a promising alternative, networks trained using such loss functions tend to introduce reconstruction artifacts that are unacceptable in most cases. In this paper, we explore how to effectively balance pixel-wise and perceptual image similarity when training a network for multi-frame super-resolution. Through extensive experimental validation, including a mean opinion score survey, we demonstrate that the proposed approach improves the reconstruction quality and we investigate its robustness against distortions in the training set, inherent to real-world cases.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Toward Exploiting Perceptual Image Similarity for Multi-frame Super-Resolution

  • Grzegorz Brzeczek,
  • Maciej Zyrek,
  • Michal Kawulok

摘要

Multi-frame super-resolution is a process of combining multiple low-resolution observations of the same scene to reconstruct a high-resolution image. State-of-the-art techniques leverage deep networks that are trained in a supervised way to increase the similarity between the reconstruction outcome and a high-resolution reference image. The similarity is typically quantified using L1 or L2 norm when computing the loss function during training. However, these pixel-wise metrics often fail to provide adequate guidance, especially for real-world training datasets with low- and high-resolution images captured using different sensors. While exploiting perceptual similarity underpinned with deep features appears to be a promising alternative, networks trained using such loss functions tend to introduce reconstruction artifacts that are unacceptable in most cases. In this paper, we explore how to effectively balance pixel-wise and perceptual image similarity when training a network for multi-frame super-resolution. Through extensive experimental validation, including a mean opinion score survey, we demonstrate that the proposed approach improves the reconstruction quality and we investigate its robustness against distortions in the training set, inherent to real-world cases.