An Effective Approach to Text Detection and Recognition in Degraded Historical Documents
摘要
In this paper, we present a method for improving text detection and recognition in historical and degraded document images by integrating Document Layout Analysis (DLA) with advanced Image Enhancement (IE) techniques and Optical Character Recognition (OCR). Our method utilizes the state-of-the-art Vision Grid Transformer for precise detection of text regions, followed by a series of enhancement processes including skew correction, Gaussian blur, binarization, and mathematical morphology to enhance the clarity and readability of the text. Evaluated on a custom dataset of scanned historical books from Project Gutenberg, our approach demonstrates superior performance in terms of Word Error Rate (WER) compared to existing methods, highlighting its effectiveness in accurately recovering text from complex and degraded document images.