<p>This paper proposes a multi-level feature fusion and high-fidelity information transfer algorithm (MFIT) to enhance the quality of mural inpainting. This method incorporates a Feature Transfer Module (FTM) designed to optimize the information flow, thereby facilitating the seamless integration of high-level semantic features with the complex low-level structural details. Moreover, we introduce gated convolutions in conjunction with recursive mechanisms in a parallel architecture to form a Recursive Gated Network (RGNet), which enables arbitrary-order spatial information exchange across multi-scale features. Finally, we integrate a Mask-Aware Pixel-Shuffle Down-Sampling Module (MPD) into an unquantized Transformer network, which effectively preserves critical information in occluded regions by exploiting spatial context within masked areas. Experimental results reveal that, compared to the baseline model, the proposed approach on the Dunhuang mural dataset achieves improvements in PSNR by 1.40% ~ 5.49%, SSIM by 0.43% ~ 1.65%, a reduction in L1 by 6.25% ~ 15.63%, and a decrease in LPIPS by 2.47% ~ 17.46%.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-level feature fusion and high-fidelity information transfer algorithm for mural inpainting

  • Zhongmin Liu,
  • Zhenbin Zhang,
  • Wenjin Hu

摘要

This paper proposes a multi-level feature fusion and high-fidelity information transfer algorithm (MFIT) to enhance the quality of mural inpainting. This method incorporates a Feature Transfer Module (FTM) designed to optimize the information flow, thereby facilitating the seamless integration of high-level semantic features with the complex low-level structural details. Moreover, we introduce gated convolutions in conjunction with recursive mechanisms in a parallel architecture to form a Recursive Gated Network (RGNet), which enables arbitrary-order spatial information exchange across multi-scale features. Finally, we integrate a Mask-Aware Pixel-Shuffle Down-Sampling Module (MPD) into an unquantized Transformer network, which effectively preserves critical information in occluded regions by exploiting spatial context within masked areas. Experimental results reveal that, compared to the baseline model, the proposed approach on the Dunhuang mural dataset achieves improvements in PSNR by 1.40% ~ 5.49%, SSIM by 0.43% ~ 1.65%, a reduction in L1 by 6.25% ~ 15.63%, and a decrease in LPIPS by 2.47% ~ 17.46%.