错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Joint Optimization of Autoencoder-Guided Attention Deep Back-Projection Network and Transformer for Document Image Enhancement and Recognition

  • Ankit Shukla,
  • Avinash Upadhyay,
  • Manoj Sharma

摘要

Accurate recognition of text from noisy, degraded, and low-resolution images has continued to be a hard problem even after years of research in computer vision. Most of the existing frameworks focus on clean high-resolution images for recognizing text, and the few recent works around degraded low-resolution images are still struggling around improving recognition at the cost of image enhancement or vice-versa. To address this, we present a framework which can be optimized jointly on enhancement and recognition loss. Toward this goal, the first contribution is the image enhancement module: an autoencoder-guided deep back-projection network that simultaneously learns denoising and super-resolution as two loss functions for image enhancement. The second contribution of this work is the use of transformer architecture in recognition module along with the image enhancement module and performs joint optimization of both modules to make the entire framework robust and resource-efficient. The transformer network is used for text recognition, and the gradient of recognition error is also backpropagated to update the weights of the enhancement module. To further improve the performance of proposed architecture on degradation and low-resolution images, BERT is used as post-processing. Experimental results elucidate the efficacy of the proposed framework and prove the robustness of the enhancement module and resource efficiency. The ablation study is also conducted to highlight the importance of each module for the improved performance.