Comparison of Existing Versus New Model of Tesseract OCR for the Gujarati Language
摘要
Gujarati is a Devanagari script that has made a substantial literary and interpersonal cultural contribution to India. The literary works of this language, which are rich in text, might disappear in the future due to aging. An optical character recognition engine is the technological solution used to preserve them. Images based on the Gujarati language can be converted to equivalent text outputs using the Tesseract OCR engine. A significant obstacle to producing text outputs of higher quality is obtaining high-resolution, noise-free images. The existing Gujarati model faces issues while identifying text written in old-type font face or typewriter-like fonts, popularly prevalent in 1960–1990. The researchers have compared the outcomes of the Tesseract OCR engine for the current language model and their language model for the Gujarati language and have submitted the attained results.