Classification of Text and Non-text Components Present in Offline Unconstrained Handwritten Documents Using Convolutional Neural Network
摘要
Identification of text parts and non-text parts present in offline unconstrained handwritten manuscripts is an essential step toward the construction of an effective optical character recognition (OCR) system. To address the said issue researchers mostly extracted handcrafted features which capture the texture information in order to recognize text or non-text components separately. In presence of noise, these types of feature descriptors badly suffer. Therefore, in this paper, a Convolutional Neural Network (CNN) is designed to separate these extracted components. To evaluate the developed model, an in-house dataset of 150 pages is created. In this dataset, the present model has achieved 85.07% accuracy. The performance of the present model is compared with three recent works where it has outperformed these existing works.