Optical character recognition plays a critical in the process of document preservation as it makes it possible to convert non-digital information such as historical texts into a digital form. However, there are numerous challenges that arises when it comes to the restoration and digitization of historic texts. These challenges include deterioration of old texts, dissimilar fonts, and lack of an annotated datasets which is often in limited volume. This paper focuses on improving the accuracy of OCR for historical documents by utilizing traditional techniques like Tesseract OCR with convolutional neural networks which is further enhanced with the synthetic data using generative artificial intelligence. Tesseract as a traditional OCR system does not work well with complex texts because of rule-based techniques. Therefore, we propose a method that uses CNNs, which allow for learning and generalization of features over multiple font styles, formats, and enhancing the recognition of historic scripts with different typologies. This research also addresses the issue of scarcity of data which is a major issue when it comes to historical document preservation. Generative AI is harnessed to create artificial datasets that mimic real. These synthetic datasets help to overcome the above restriction by increasing the volume of training data, which in turn enhances the OCR systems in recognizing even more types of scripts and degraded texts. The relative performance of the proposed hybrid approach is measured against standard OCR benchmarks which reveals improvement in performance metrics.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing OCR for Historical Documents Using CNNs and GenAI

  • Rizul Bhardwaj,
  • Sneha Kumari,
  • Joelin J. Jacob,
  • Anchal Kamra

摘要

Optical character recognition plays a critical in the process of document preservation as it makes it possible to convert non-digital information such as historical texts into a digital form. However, there are numerous challenges that arises when it comes to the restoration and digitization of historic texts. These challenges include deterioration of old texts, dissimilar fonts, and lack of an annotated datasets which is often in limited volume. This paper focuses on improving the accuracy of OCR for historical documents by utilizing traditional techniques like Tesseract OCR with convolutional neural networks which is further enhanced with the synthetic data using generative artificial intelligence. Tesseract as a traditional OCR system does not work well with complex texts because of rule-based techniques. Therefore, we propose a method that uses CNNs, which allow for learning and generalization of features over multiple font styles, formats, and enhancing the recognition of historic scripts with different typologies. This research also addresses the issue of scarcity of data which is a major issue when it comes to historical document preservation. Generative AI is harnessed to create artificial datasets that mimic real. These synthetic datasets help to overcome the above restriction by increasing the volume of training data, which in turn enhances the OCR systems in recognizing even more types of scripts and degraded texts. The relative performance of the proposed hybrid approach is measured against standard OCR benchmarks which reveals improvement in performance metrics.