GujFormer: A Vision Transformer-Based Architecture for Gujarati Handwritten Character Recognition
摘要
It is challenging and requires automatic recognition of handwritten letters and numbers, especially in the era of digitalization. However, several literatures exist in major and nationalized languages. In the last few years, in an era of the paperless office, it has been necessary to cover all the most utilized languages in the region. Optical character recognition (OCR) with deep learning has achieved remarkable performance in various character recognition and interpretation tasks. In this study, the authors proposed a vision transformer (ViT) base approach (GujFormer) to recognize Gujarati handwritten characters. The proposed study used a multihead self-attention model to enhance feature learning. The proposed methodology used three different datasets having handwritten characters, numbers, and printed Gujarati characters. The vision transformer achieved 98.31% accuracy for handwritten digits and 97.92% accuracy for handwritten characters. The methodology used multilayer perceptron (MLP) as a classification layer. Simulation of the proposed methodology was also compared with other deep learning CNN models like VGG16, InceptionV3, DenseNet121, and others. Authors have also analyzed recognition variation over different self-attention heads. The proposed methodology has achieved remarkable performance over vision transformer for Gujarat handwritten character recognition.