A structural information-guided cross-modal method for damaged inscription inpainting via vision-language models
摘要
Restoring inscriptions is crucial for preserving cultural heritage. Current methods primarily focus on visual-level generation and inpainting, ignoring glyph structure information. However, the structural integrity of Chinese characters is frequently compromised in damaged inscription images. To address this challenge, we propose a structural information-guided cross-modal inpainting method. Our dual-branch network includes an inpainting branch and a structure branch. Firstly, to compensate for missing structural information, we pretrain a vision-language model to obtain high-quality glyph structure representations by decomposing each Chinese character into components and structural relationships. Secondly, the glyph structure representation guides the structure branch to optimize features from the damaged character image, producing features that contain more glyph structure information. Thirdly, a feature interaction mechanism injects the optimized features into the inpainting branch, and an adaptive style embedding module improves restoration accuracy in style, structure, and detail. Moreover, a feature sharing module alleviates potential conflicts between branches.