Classification and feature extraction of text from hindi document for optical character recognition
摘要
In the area of optical character recognition, Hindi documents have a complex research topic. Much research has been done into OCR techniques with different scripts, including Japanese, Roman, Korean, Chinese, and some Indian languages. The amount of OCR study done on the Devanagari script is small. The Challenge of character recognition can be separated into two sections: handwritten characters and printed character recognition from documents. This paper presents a relevant surveys and compares different methods used in the character’s recognition, classification, and feature extraction. This comparison and analysis will show a technical review of other parameters and evaluation of different classifiers used in various existing techniques. This extensive study shows how multiple classification algorithms and feature extraction for offline Devanagari character recognition work. Several concerns and obstacles relating to recognizing Indian scripts are examined, indicating possible future study directions. It has been concluded that hybrid feature extraction and classification approaches can provide the most accurate findings. We have compared different methodologies regarding feature extraction strategies, categorization, and accuracy of various researchers.