SemanticAvatar: human surface reconstruction based on semantically consistent biplane features
摘要
Efficient semantic embedding is important for fine 3D human surfaces or avatars reconstructed from single images. However, this embedding is weak among existing methods. In particular, the widely adopted triplane features extracted from single images often exhibit the semantic inconsistencies and mutual interferences. This paper aims at resolving those limitations for better single-image-based reconstruction. First, simplification of the triplane features by semantically consistent two-plane representation is introduced. Then, based on the biplane features, a novel diffusion-based framework, SemanticAvatar, is proposed. It adopts the large visual models for inferring the occluded view as the complement and further obtains semantically consistent biplane features through shared inference. Two feature-level semantic enhancement strategies are further incorporated: (1) a semantic alignment module which embeds the image semantics by multi-scale aligning image features with the latent representations, thus promoting semantic consistency during the diffusion process; (2) a human-structure-based feature cropping strategy, which enables more precise semantic alignment by body-part-based decomposition, ultimately yielding more accurate reconstruction details. Experimental results demonstrate the effectiveness of SemanticAvatar.