Image super resolution of Thangka murals using multi-scale feature assisted transformer and hybrid attention
摘要
Thangka murals are vital Tibetan cultural heritage. However, existing digital images face insufficient clarity, often limited to resolutions below 1024 × 1024 pixels, which hinders cultural preservation and analysis. Despite their potential, Transformer architectures encounter two key bottlenecks in Thangka reconstruction: first, single-scale self-attention mechanisms struggle to represent multi-scale artistic features; second, existing methods insufficiently utilize global information during reconstruction. To address these issues, this study proposes a Thangka super-resolution model integrating Multi-Scale Feature Assisted Transformer (MSFA-Transformer) and Hybrid Attention Block (HAB). MSFA-Transformer introduces a parallel multi-scale feature modulation branch to enhance scale-aware representation beyond window-based self-attention. HAB adopts multi-dimensional attention fusion, combining channel attention with spatial attention to expand information utilization. On 1024 × 1024 Thangka dataset, our ×2 super-resolution achieves 34.47 dB PSNR, surpassing CNN-based RCAN by 0.26 dB and Transformer-based SwinIR by 0.18 dB, demonstrating superior restoration of intricate patterns and natural color transitions.