In this paper, we present a method for improving text detection and recognition in historical and degraded document images by integrating Document Layout Analysis (DLA) with advanced Image Enhancement (IE) techniques and Optical Character Recognition (OCR). Our method utilizes the state-of-the-art Vision Grid Transformer for precise detection of text regions, followed by a series of enhancement processes including skew correction, Gaussian blur, binarization, and mathematical morphology to enhance the clarity and readability of the text. Evaluated on a custom dataset of scanned historical books from Project Gutenberg, our approach demonstrates superior performance in terms of Word Error Rate (WER) compared to existing methods, highlighting its effectiveness in accurately recovering text from complex and degraded document images.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Effective Approach to Text Detection and Recognition in Degraded Historical Documents

  • Percy Maldonado-Quispe,
  • Helio Pedrini

摘要

In this paper, we present a method for improving text detection and recognition in historical and degraded document images by integrating Document Layout Analysis (DLA) with advanced Image Enhancement (IE) techniques and Optical Character Recognition (OCR). Our method utilizes the state-of-the-art Vision Grid Transformer for precise detection of text regions, followed by a series of enhancement processes including skew correction, Gaussian blur, binarization, and mathematical morphology to enhance the clarity and readability of the text. Evaluated on a custom dataset of scanned historical books from Project Gutenberg, our approach demonstrates superior performance in terms of Word Error Rate (WER) compared to existing methods, highlighting its effectiveness in accurately recovering text from complex and degraded document images.