Word-Diffusion: Diffusion-Based Handwritten Text Word Image Generation
摘要
Generating realistic handwritten word images that closely resemble a target style remains a challenging task in document image analysis. In recent years, deep learning techniques, such as Latent Diffusion Models (LDM), have shown promise in generating styled handwritten text. However, these models face significant challenges when creating images for ‘Out of Vocabulary’ (OOV) words, impacting their overall effectiveness. In this paper, we introduce an extended diffusion-based Handwritten generation method that incorporates a novel conditioning mechanism. It is based on the Pyramidal Histogram of Shapes (PHOS) representation, which takes into account the spatial and structural characteristics of the target handwriting style. By conditioning the diffusion model on input text, PHOS vector, and writer ID, our approach enables the generation of handwritten word images. Notably, our approach outperforms the original diffusion model, which only uses text and writer ID as conditions, in generating both in-sample and out-of-sample. Furthermore, we have developed a faster inference method that significantly reduces the number of steps required for generating the output. Through qualitative and quantitative evaluations, we demonstrate the effectiveness of our proposed method.