错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

GujFormer: A Vision Transformer-Based Architecture for Gujarati Handwritten Character Recognition

  • Deep R. Kothadiya,
  • Chintan Bhatt,
  • Aayushi Chaudhari,
  • Nilkumar Sinojiya

摘要

It is challenging and requires automatic recognition of handwritten letters and numbers, especially in the era of digitalization. However, several literatures exist in major and nationalized languages. In the last few years, in an era of the paperless office, it has been necessary to cover all the most utilized languages in the region. Optical character recognition (OCR) with deep learning has achieved remarkable performance in various character recognition and interpretation tasks. In this study, the authors proposed a vision transformer (ViT) base approach (GujFormer) to recognize Gujarati handwritten characters. The proposed study used a multihead self-attention model to enhance feature learning. The proposed methodology used three different datasets having handwritten characters, numbers, and printed Gujarati characters. The vision transformer achieved 98.31% accuracy for handwritten digits and 97.92% accuracy for handwritten characters. The methodology used multilayer perceptron (MLP) as a classification layer. Simulation of the proposed methodology was also compared with other deep learning CNN models like VGG16, InceptionV3, DenseNet121, and others. Authors have also analyzed recognition variation over different self-attention heads. The proposed methodology has achieved remarkable performance over vision transformer for Gujarat handwritten character recognition.