<p>Handwriting recognition is a challenging task, especially for national languages, and even regional one. Many scholars have spent years researching optical character recognition (OCR) and contributing to developing script-specific OCR systems. The Assamese script consists of 10 digits, 11 vowels, 41 consonants and around 223 compound characters. While there are existing Optical Character Recognition (OCR) systems for recognizing Assamese basic characters, no research focuses on handwritten Assamese compound characters. This challenge arises from the script’s large number of character classes, the visual similarity between certain character pairs, and the high variability in individual handwriting styles. Additionally, there is no publicly available dataset dedicated to Assamese handwritten compound characters. To address this gap, we proposed a benchmark dataset for unconstrained Assamese handwritten compound characters used in the standard Assamese literature. A thorough survey shows that at least 223 compound characters exist in the Assamese script. In this work, all 223 compound characters had been considered. The dataset is built from contributors of various age groups and occupations, enhancing its diversity and potential for generalization across real world scenarios. Altogether, 135,584 isolated images belonging to 223 Assamese compound character classes have been collected. The dataset was split into three subsets: 70% for training, 20% for validation, and 10% for testing. A stratified random split is used to ensure that each subset maintains the same class distribution as the overall dataset. To see the robustness of the developed dataset; different deep learning models namely InceptionV3, ResNet50, EfficientNetB0, MobileNet and a customized CNN have been used. Our proposed modified CNN achieved a recognition accuracy of 99.07% on the developed dataset, with reduced convergence time compared to other evaluated models. The dataset is available at: <a href="https://figshare.com/s/9ba2b2f4d12e296079bb">https://figshare.com/s/9ba2b2f4d12e296079bb</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A benchmark image dataset of unconstrained isolated Assamese handwritten compound characters

  • Gaurab Khakhlari,
  • Sanjib Kr. Kalita,
  • Minakshi Gogoi,
  • Rupjyoti Ray,
  • Shikhar Kr. Sarma

摘要

Handwriting recognition is a challenging task, especially for national languages, and even regional one. Many scholars have spent years researching optical character recognition (OCR) and contributing to developing script-specific OCR systems. The Assamese script consists of 10 digits, 11 vowels, 41 consonants and around 223 compound characters. While there are existing Optical Character Recognition (OCR) systems for recognizing Assamese basic characters, no research focuses on handwritten Assamese compound characters. This challenge arises from the script’s large number of character classes, the visual similarity between certain character pairs, and the high variability in individual handwriting styles. Additionally, there is no publicly available dataset dedicated to Assamese handwritten compound characters. To address this gap, we proposed a benchmark dataset for unconstrained Assamese handwritten compound characters used in the standard Assamese literature. A thorough survey shows that at least 223 compound characters exist in the Assamese script. In this work, all 223 compound characters had been considered. The dataset is built from contributors of various age groups and occupations, enhancing its diversity and potential for generalization across real world scenarios. Altogether, 135,584 isolated images belonging to 223 Assamese compound character classes have been collected. The dataset was split into three subsets: 70% for training, 20% for validation, and 10% for testing. A stratified random split is used to ensure that each subset maintains the same class distribution as the overall dataset. To see the robustness of the developed dataset; different deep learning models namely InceptionV3, ResNet50, EfficientNetB0, MobileNet and a customized CNN have been used. Our proposed modified CNN achieved a recognition accuracy of 99.07% on the developed dataset, with reduced convergence time compared to other evaluated models. The dataset is available at: https://figshare.com/s/9ba2b2f4d12e296079bb.