Emotion guided multi scale feature fusion for Qinqiang Opera video restoration
摘要
As an important form of China’s intangible cultural heritage, Qinqiang Opera relies on precise facial expression to convey artistic meaning. Conventional video restoration methods often neglect emotion-sensitive regions, leading to loss of expressive detail. We propose an emotion-guided restoration framework that extracts temporal emotion tone vectors from audio using an enhanced MHAtt_ResNet model, and integrates them into the restoration stage through an Emotion-Gated Module and Cross-Stage Feature Fusion. Further, an Emotion-Driven Temporal Aggregation Module with a denoising subnetwork restores facial dynamics under noise. Experiments on Qinqiang videos show that our method achieves superior PSNR, SSIM, MOS, as well as lower temporal metrics (tOF, TGMSD), compared with EDVR, BasicVSR + +, FastDVDnet, UHDVD, and TACE. The model restores subtle facial details while maintaining smooth temporal coherence, significantly improving artistic fidelity. This work provides a technically advanced and culturally adaptive solution for digital preservation of performance heritage.