<p>Light field (LF) can capture both the intensity and direction information of light rays and is widely used in applications such as depth estimation. The emergence of LF cameras has promoted the dissemination of LF media, but the captured LF images suffer from low spatial resolution while providing high angular sampling. Besides, the current mainstream media is still dominated by 2D images. Consequently, reconstructing high-resolution 2D images into dense LF images is an effective way to further promote the development of LFs. To this end, this paper proposes an LF synthesis method based on monocular images that integrate implicit and explicit depth information. Specifically, the proposed method adopts a complementary dual-branch structure, i.e., an implicit depth branch and an explicit depth branch, to handle large parallax and recover detailed textures. First, the former leverages a Swin Transformer-based network to explore long-range spatial dependencies to implicitly reconstruct the target view. Meanwhile, a position-aware feature fusion module is designed in the network to effectively embed angular priors into content features. Secondly, the latter employs a multiplane image representation and combines the inferred depth map to reconstruct the target view. To facilitate the generation of the multiplane image model, a local–global feature extraction module is designed to simultaneously capture texture and geometric information. Finally, a mask-guided fusion module is constructed to adaptively fuse the reconstruction results of the two branches. Experimental results show that the proposed method is superior to the existing representative methods in both quantitative and qualitative comparisons. Simply put, the proposed method increases PSNR by 5.475% and SSIM by 2.599%, which not only enhances visual perception but also facilitates subsequent LF applications. In addition, detailed ablation studies validate the contribution of core components in the proposed method.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Monocular Image-Based Light Field Synthesis by Integrating Implicit and Explicit Depth Information

  • Mingxing Fu,
  • Yeyao Chen,
  • Chongchong Jin,
  • Zongju Peng,
  • Haiyong Xu,
  • Gangyi Jiang

摘要

Light field (LF) can capture both the intensity and direction information of light rays and is widely used in applications such as depth estimation. The emergence of LF cameras has promoted the dissemination of LF media, but the captured LF images suffer from low spatial resolution while providing high angular sampling. Besides, the current mainstream media is still dominated by 2D images. Consequently, reconstructing high-resolution 2D images into dense LF images is an effective way to further promote the development of LFs. To this end, this paper proposes an LF synthesis method based on monocular images that integrate implicit and explicit depth information. Specifically, the proposed method adopts a complementary dual-branch structure, i.e., an implicit depth branch and an explicit depth branch, to handle large parallax and recover detailed textures. First, the former leverages a Swin Transformer-based network to explore long-range spatial dependencies to implicitly reconstruct the target view. Meanwhile, a position-aware feature fusion module is designed in the network to effectively embed angular priors into content features. Secondly, the latter employs a multiplane image representation and combines the inferred depth map to reconstruct the target view. To facilitate the generation of the multiplane image model, a local–global feature extraction module is designed to simultaneously capture texture and geometric information. Finally, a mask-guided fusion module is constructed to adaptively fuse the reconstruction results of the two branches. Experimental results show that the proposed method is superior to the existing representative methods in both quantitative and qualitative comparisons. Simply put, the proposed method increases PSNR by 5.475% and SSIM by 2.599%, which not only enhances visual perception but also facilitates subsequent LF applications. In addition, detailed ablation studies validate the contribution of core components in the proposed method.