Sketch-text-driven rectified flow for identity-preserving 4D face generation
摘要
The limitations of single-modality control in preserving facial identity and describing temporal expressions have motivated sketch-text dual-driven 4D face generation, which provides a flexible paradigm for dynamic digital face synthesis in applications requiring precise identity customization and controllable expression manipulation. However, this task remains challenging due to the synthetic-to-real domain gap in sparse sketches, cross-modal interference between heterogeneous conditions, and the scarcity of paired sketch-text 4D mesh data. To address these challenges, we propose sketch-text-driven rectified flow (STDRF), a conditional rectified-flow framework for identity-preserving and semantically controllable 4D face generation. STDRF learns a velocity field that transports Gaussian noise to the target 4D facial motion manifold under sketch-structural and text-semantic dual guidance. To obtain robust identity priors from sparse sketches, we design a sketch encoder enhanced by Geometric Contour and Texture Detail (GCTD) preprocessing and MixStyle domain adaptation. To reduce cross-modal interference, a dual-path independent cross-attention module based on IP-Adapter is designed to inject sketch features and text semantics into an Attention DiffusionNet Block (ADNB)-based denoising backbone in parallel. Furthermore, a multimodal triplet dataset is constructed by pairing 4D facial mesh sequences with synthetic sketches and hierarchical text descriptions. On the subject-independent test split, STDRF achieves an MVE of 0.81