错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comprehensive Dataset Building of Isolated Handwritten Sanskrit Characters

  • G. Dhruva,
  • Vrinda Kore,
  • M. Vijitha,
  • Sahana Rao,
  • P. Preethi

摘要

Developing character recognition systems for Indic scripts is both challenging and promising for researchers. This complexity arises from the intricate nuances among characters, the diverse writing styles, a vast range of classes, and the absence of a comprehensive dataset. The absence of such a dataset for Sanskrit characters has hindered the advancement of sophisticated computational character recognition systems. Notably, the Sanskrit script serves as the foundation for various Indic languages like Hindi, Bangla, and Gujarati. The establishment of benchmark datasets could pave the way for developing scalable character recognition systems. This paper concentrates on crafting an extensive dataset for isolated handwritten Sanskrit characters, focusing on recognizing Sanskrit vowels, consonants, modifiers, and compound characters. The dataset was meticulously gathered from contributors across diverse ages, genders, qualifications, and professions. Subsequently, the dataset was digitized and underwent extensive pre-processing by utilizing image processing techniques, and characters were categorized into specific classes through machine learning methods. The primary goal is to enable the identification of handwritten Sanskrit characters, particularly emphasizing the recognition of compound characters and modifiers, aspects often overlooked in existing research. The comprehensive dataset will be made available online, serving as an invaluable resource for researchers working on efficient and scalable character recognition frameworks in Sanskrit or its derived languages.