错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

High-definition multi-scale voice-driven facial animation: enhancing lip-sync clarity and image detail

  • Long Zhang,
  • QingHua Zhou,
  • Shuai Tang,
  • Yunxiang Chen

摘要

The advancement of voice-driven facial video generation networks has significantly enriched the realm of AI-generated video content. However, generated facial videos often suffer from blurred mouth regions, compromising their visual authenticity. In this study, we propose HighDefWav2lip-MS, an extension of the Wav2lip algorithm, aimed at enhancing facial video clarity, particularly in the mouth region. We achieve this by expanding the input size from 96 × 96 to 384 × 384 pixels, optimizing the architecture, and introducing a multi-scale structural similarity (MS-SSIM) loss along with an L2 loss into the generator. This modification not only refines lip synchronization but also improves image detail retention. We evaluate our approach using a variety of metrics, including FID, PSNR, IS, and LPIPS. Compared to existing methods, our results show superior performance, with a significant reduction in FID scores of 61.2% and an increase in PSNR values of 4.29%. Additionally, our method significantly improves IS by 3.32% and significantly reduces LPIPS by 23.1%, indicating enhanced perceptual quality and realism. These results underscore the effectiveness of HighDefWav2Lip-MS in generating high-resolution facial videos with superior clarity and realism.