<p>Tibetan painting is an important form of visual cultural heritage, characterized by complex iconographic structures, dense decorative patterns, and distinctive stylistic conventions. Generating Tibetan painting images from sparse sketches and textual descriptions remains challenging because sparse sketches provide incomplete structural cues, while culturally plausible synthesis requires semantic consistency, structural coherence, and style-aware representation. To address these challenges, we propose STP-Diff, a sketch- and text-guided diffusion framework for Tibetan painting generation. STP-Diff integrates sketch semantic parsing, sparse-to-dense structure enhancement, multi-condition semantic fusion, and line-conditioned style prior learning to guide diffusion with structural and semantic constraints. We further construct an extended HHTP dataset with paired sketch, line drawing, color image, and text annotations. Experiments on HHTP and SketchyCOCO demonstrate that STP-Diff improves structural preservation, semantic alignment, visual plausibility, and generalization, providing a computational approach for digital preservation and creative reuse of Tibetan painting heritage.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A sketch-and-text guided diffusion framework for Tibetan painting generation for digital heritage preservation

  • Fubo Wang,
  • Mingcong Dang,
  • Zeyu Jia,
  • Wanyi Zhao,
  • Shengling Geng

摘要

Tibetan painting is an important form of visual cultural heritage, characterized by complex iconographic structures, dense decorative patterns, and distinctive stylistic conventions. Generating Tibetan painting images from sparse sketches and textual descriptions remains challenging because sparse sketches provide incomplete structural cues, while culturally plausible synthesis requires semantic consistency, structural coherence, and style-aware representation. To address these challenges, we propose STP-Diff, a sketch- and text-guided diffusion framework for Tibetan painting generation. STP-Diff integrates sketch semantic parsing, sparse-to-dense structure enhancement, multi-condition semantic fusion, and line-conditioned style prior learning to guide diffusion with structural and semantic constraints. We further construct an extended HHTP dataset with paired sketch, line drawing, color image, and text annotations. Experiments on HHTP and SketchyCOCO demonstrate that STP-Diff improves structural preservation, semantic alignment, visual plausibility, and generalization, providing a computational approach for digital preservation and creative reuse of Tibetan painting heritage.