Challenges in Making OCR of Gujarati Newspaper
摘要
The newspaper is a great source of information. We used to receive news about daily affairs, sports, general issues, religion, criminals, and other topics. It is difficult to collect and keep newspapers as hard copies for potential use to find information about the past. Therefore, we must keep newspapers in digital format. In order to digitize a newspaper image, it must be scanned, segmented into pages, images, and characters, then recognized and saved as a text file. Numerous studies have been conducted to identify text on printed documents at the worldwide or national level. Yet there haven’t been many studies done to identify Gujarati newspaper text. We have discussed a number of issues that may arise when creating OCR for a Gujarati newspaper in this paper.