Audio-driven facial animation in the wild using a multi-style identity preservation framework
摘要
Audio-driven facial animation aims to generate talking-face videos that are synchronized with speech while preserving the visual identity of a target subject. Although recent methods have achieved strong lip-audio synchronization, maintaining identity consistency and temporal stability under unconstrained in-the-wild conditions remains challenging. Variations in illumination, pose, image quality, occlusion, and audio noise can lead to identity drift, frame-to-frame jitter, and degraded visual realism. This paper proposes the Multi-Style Identity Preservation Framework (MS-IPF), a dual-branch architecture that separates phoneme-driven motion representation from identity-related appearance representation. The proposed Style Adapter Fusion Module injects multi-scale identity styles into the motion decoder through feature-wise affine modulation, allowing the model to preserve both coarse facial structure and fine appearance details during generation. To support evaluation under realistic conditions, we introduce WildFaces-AudioSync, an identity-disjoint audiovisual dataset designed for assessing identity preservation, lip-audio synchronization, temporal stability, and robustness in heterogeneous recording environments. Experimental results show that MS-IPF improves identity preservation, synchronization quality, visual realism, and temporal coherence compared with representative baseline methods, while maintaining single-pass inference. Additional analyses highlight the effects of fast speech, noisy audio, language imbalance, source-image quality, and computational overhead.