<p>As an important form of China’s intangible cultural heritage, Qinqiang Opera relies on precise facial expression to convey artistic meaning. Conventional video restoration methods often neglect emotion-sensitive regions, leading to loss of expressive detail. We propose an emotion-guided restoration framework that extracts temporal emotion tone vectors from audio using an enhanced MHAtt_ResNet model, and integrates them into the restoration stage through an Emotion-Gated Module and Cross-Stage Feature Fusion. Further, an Emotion-Driven Temporal Aggregation Module with a denoising subnetwork restores facial dynamics under noise. Experiments on Qinqiang videos show that our method achieves superior PSNR, SSIM, MOS, as well as lower temporal metrics (tOF, TGMSD), compared with EDVR, BasicVSR + +, FastDVDnet, UHDVD, and TACE. The model restores subtle facial details while maintaining smooth temporal coherence, significantly improving artistic fidelity. This work provides a technically advanced and culturally adaptive solution for digital preservation of performance heritage.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Emotion guided multi scale feature fusion for Qinqiang Opera video restoration

  • Tian Xia,
  • Yili Yuan,
  • Binzhi Qi,
  • Xiaojun Liu,
  • Kaiyue Yang

摘要

As an important form of China’s intangible cultural heritage, Qinqiang Opera relies on precise facial expression to convey artistic meaning. Conventional video restoration methods often neglect emotion-sensitive regions, leading to loss of expressive detail. We propose an emotion-guided restoration framework that extracts temporal emotion tone vectors from audio using an enhanced MHAtt_ResNet model, and integrates them into the restoration stage through an Emotion-Gated Module and Cross-Stage Feature Fusion. Further, an Emotion-Driven Temporal Aggregation Module with a denoising subnetwork restores facial dynamics under noise. Experiments on Qinqiang videos show that our method achieves superior PSNR, SSIM, MOS, as well as lower temporal metrics (tOF, TGMSD), compared with EDVR, BasicVSR + +, FastDVDnet, UHDVD, and TACE. The model restores subtle facial details while maintaining smooth temporal coherence, significantly improving artistic fidelity. This work provides a technically advanced and culturally adaptive solution for digital preservation of performance heritage.