<p>Ancient manuscripts require digitization to combat deterioration and enable preservation. This study addresses challenges in multilingual document processing, including noise, texture variations, and ink fading, by proposing GLADUnet – a Global-Local Attention Dynamic UNet for denoising and binarization. The method employs convolutional neural networks (CNNs) with Spatial Group-wise Enhancement to extract character features, integrates a global-local attention module for noise reduction, and utilizes a UNet decoder with dynamic upsampling for precise image reconstruction. We construct a multilingual dataset covering 12 scripts to evaluate the framework. The experimental results demonstrate that GLADUnet shows excellent performance and can effectively handle the noise reduction and binarization tasks of multilingual ancient book document images. The model’s cross-dataset validation confirms its robustness for historical document processing tasks. This approach significantly improves optical character recognition (OCR) accuracy by effectively separating text from complex backgrounds and noise.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

GLADUnet: global-local attention and dynamic upsampling for ancient manuscript denoising and binarization

  • Hai Guo,
  • Dawei Zhu,
  • Lingling Tong,
  • Jingying Zhao,
  • Aoqi Jiang

摘要

Ancient manuscripts require digitization to combat deterioration and enable preservation. This study addresses challenges in multilingual document processing, including noise, texture variations, and ink fading, by proposing GLADUnet – a Global-Local Attention Dynamic UNet for denoising and binarization. The method employs convolutional neural networks (CNNs) with Spatial Group-wise Enhancement to extract character features, integrates a global-local attention module for noise reduction, and utilizes a UNet decoder with dynamic upsampling for precise image reconstruction. We construct a multilingual dataset covering 12 scripts to evaluate the framework. The experimental results demonstrate that GLADUnet shows excellent performance and can effectively handle the noise reduction and binarization tasks of multilingual ancient book document images. The model’s cross-dataset validation confirms its robustness for historical document processing tasks. This approach significantly improves optical character recognition (OCR) accuracy by effectively separating text from complex backgrounds and noise.