<p>Face super-resolution (FSR) techniques play a pivotal role in enhancing the resolution and quality of low-resolution facial images, which is crucial for applications such as surveillance, forensic analysis, and face recognition. Traditional convolution- and MLP-based models often struggle to effectively capture facial structural information and model complex nonlinear relationships, resulting in suboptimal restoration of fine details. To address these limitations, this paper introduces a novel FSR method based on iterative collaboration of KAN–MLP hybrid attention and landmark estimation (ICKL). The proposed method incorporates a spatial-frequency multi-scale block (SFMSB) to extract multi-scale features and integrate frequency-domain information, thereby enhancing the capture of key facial structural details. Furthermore, a KAN–MLP hybrid attention group (KHAG) is designed to refine the extracted features, combining Kolmogorov–Arnold network-based window attention blocks (KAN–WAB) and MLP-based window attention blocks (MLP–WAB) to improve the modeling of complex nonlinear relationships. Through an iterative collaboration mechanism between the reconstruction network and the facial landmark estimation network, the proposed ICKL method achieves more accurate and consistent facial reconstructions. Extensive experiments on the CelebA and Helen datasets under standard 8<InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mo>×</mo> </math></EquationSource> </InlineEquation> and 16<InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mo>×</mo> </math></EquationSource> </InlineEquation> super-resolution settings demonstrate the superior performance of our method. Specifically, on the 8<InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mo>×</mo> </math></EquationSource> </InlineEquation> super-resolution task, ICKL achieves 27.73dB PSNR, 0.8092 SSIM, and 0.0849 LPIPS on the CelebA dataset, outperforming state-of-the-art methods in both accuracy and visual quality. Our code is available at <a href="https://github.com/sslavender/ICKL">https://github.com/sslavender/ICKL</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhanced face super-resolution via iterative collaboration of KAN–MLP hybrid attention and landmark estimation

  • Zhiyong An,
  • Changteng Shi,
  • Yongliang Wang

摘要

Face super-resolution (FSR) techniques play a pivotal role in enhancing the resolution and quality of low-resolution facial images, which is crucial for applications such as surveillance, forensic analysis, and face recognition. Traditional convolution- and MLP-based models often struggle to effectively capture facial structural information and model complex nonlinear relationships, resulting in suboptimal restoration of fine details. To address these limitations, this paper introduces a novel FSR method based on iterative collaboration of KAN–MLP hybrid attention and landmark estimation (ICKL). The proposed method incorporates a spatial-frequency multi-scale block (SFMSB) to extract multi-scale features and integrate frequency-domain information, thereby enhancing the capture of key facial structural details. Furthermore, a KAN–MLP hybrid attention group (KHAG) is designed to refine the extracted features, combining Kolmogorov–Arnold network-based window attention blocks (KAN–WAB) and MLP-based window attention blocks (MLP–WAB) to improve the modeling of complex nonlinear relationships. Through an iterative collaboration mechanism between the reconstruction network and the facial landmark estimation network, the proposed ICKL method achieves more accurate and consistent facial reconstructions. Extensive experiments on the CelebA and Helen datasets under standard 8 \(\times \) × and 16 \(\times \) × super-resolution settings demonstrate the superior performance of our method. Specifically, on the 8 \(\times \) × super-resolution task, ICKL achieves 27.73dB PSNR, 0.8092 SSIM, and 0.0849 LPIPS on the CelebA dataset, outperforming state-of-the-art methods in both accuracy and visual quality. Our code is available at https://github.com/sslavender/ICKL.