HD-Tex: Leveraging Structural Priors for High-Fidelity Texture Synthesis
摘要
High-quality texture generation for 3D meshes from textual descriptions represents a critical yet technically challenging task in computer graphics and vision. While recent advancements in optimization techniques—particularly those based on Score Distillation Sampling (SDS)—have demonstrated promising outcomes, they frequently produce textures with visual artifacts and inconsistencies, thereby limiting their practical deployment. In this paper, we propose HD-Tex, a novel text-guided 3D texturing framework that strikes an effective balance between computational efficiency and texture fidelity. Our key contributions lie in two aspects: (1) the introduction of a 3D structural prior to effectively mitigate seam artifacts by enforcing geometric consistency across views; and (2) the design of a dual-resolution self-supervised optimization strategy that enhances fine-grained texture details while preserving global structural coherence. Together, these components enable high-fidelity texture generation that is both locally detailed and globally consistent. Our approach consists of two stages: (1) a depth-aware, multimodal diffusion module generates four anchor views, which are projected into UV space to create an initial texture map; and (2) a targeted optimization routine selectively updates regions based on depth cues, followed by a refinement process operating across dual resolutions. Comprehensive experiments confirm that HD-Tex outperforms existing baselines, delivering significantly improved texture realism and consistency.