错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Working of the Tesseract OCR on Different Fonts of Gujarati Language

  • Kartik Joshi,
  • Harshal Arolkar

摘要

An optical character recognition engine is the technological solution for preserving books and manuscripts that may soon be lost due to deterioration. In digital form, documents and/or text files are editable, searchable, and shareable. To save them from getting destroyed, documents and/or text files need to be scanned/converted into digital form and passed onto the optical character recognition engine to generate the digital text file. For a large amount of data, manual typing and conversion is nearly impossible. In this paper, the authors have tried to analyze the working of the Tesseract OCR engine for the images that contain Gujarati text.