错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automatic Text Recognition from Image Dataset Using Optical Character Recognition and Deep Learning Techniques

  • Ishan Rao,
  • Prathmesh Shirgire,
  • Sanket Sanganwar,
  • Kedar Vyawhare,
  • S. R. Vispute

摘要

Optical Character Recognition (OCR) has become quite well known in the last few years, because it has applications in many sectors. In this paper, we look at the basics of OCR and discuss a few popular datasets that can help one get started with OCR. We aim to comprehensively analyze the research done on OCR with a variety of algorithms. We also look at some popular machine learning models that are used in building OCR systems. Support vector machines and convolutional neural networks are examples of these machine learning models. They have been explained briefly. We use the tesseract tool to extract text. This paper serves as a basic guide to getting started with OCR. Using the datasets and algorithms described, one can start their journey in OCR. We provide our own analysis on handwritten digit dataset using different models in machine learning and then compare their accuracy. We show the use of the popular OCR tool Tesseract by extracting data from a report. The data drawn out from the report is loaded in a database. This is a useful application that can be extended in the future to store details of different reports in medical, transportation, banking, and many other fields to create a paperless environment.