Toward Exploiting Perceptual Image Similarity for Multi-frame Super-Resolution
摘要
Multi-frame super-resolution is a process of combining multiple low-resolution observations of the same scene to reconstruct a high-resolution image. State-of-the-art techniques leverage deep networks that are trained in a supervised way to increase the similarity between the reconstruction outcome and a high-resolution reference image. The similarity is typically quantified using L1 or L2 norm when computing the loss function during training. However, these pixel-wise metrics often fail to provide adequate guidance, especially for real-world training datasets with low- and high-resolution images captured using different sensors. While exploiting perceptual similarity underpinned with deep features appears to be a promising alternative, networks trained using such loss functions tend to introduce reconstruction artifacts that are unacceptable in most cases. In this paper, we explore how to effectively balance pixel-wise and perceptual image similarity when training a network for multi-frame super-resolution. Through extensive experimental validation, including a mean opinion score survey, we demonstrate that the proposed approach improves the reconstruction quality and we investigate its robustness against distortions in the training set, inherent to real-world cases.