Optical Character Recognition
摘要
Optical character recognition, shortly called OCR, is a technology used to detect text characters from scanned text documents or digital images captured by a camera and convert them into an editable data. To put it in different terms, the OCR technology extracts machine-encoded text from the text characters recognized within scanned documents or images. OCR technology began to emerge in the early twentieth century. The primary purpose of early OCR systems was to identify printed text in documents that were typewritten or typeset. Later on, techniques like pattern matching and template matching were employed to create commercial character recognition systems. With the rise in popularity of machine learning methods as well as the development of advanced computing capabilities, these techniques gained prominence in implementing highly accurate OCR systems. Deep learning methods have improved OCR performance tremendously in the recent years, making it possible to recognize intricate handwritten text.