错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparison of Existing Versus New Model of Tesseract OCR for the Gujarati Language

  • Kartik Joshi,
  • Harshal Arolkar

摘要

Gujarati is a Devanagari script that has made a substantial literary and interpersonal cultural contribution to India. The literary works of this language, which are rich in text, might disappear in the future due to aging. An optical character recognition engine is the technological solution used to preserve them. Images based on the Gujarati language can be converted to equivalent text outputs using the Tesseract OCR engine. A significant obstacle to producing text outputs of higher quality is obtaining high-resolution, noise-free images. The existing Gujarati model faces issues while identifying text written in old-type font face or typewriter-like fonts, popularly prevalent in 1960–1990. The researchers have compared the outcomes of the Tesseract OCR engine for the current language model and their language model for the Gujarati language and have submitted the attained results.