Robust Dataset of Digital Handwritten Devanagari Script Images and Implementing Deep Learning for Handwritten Devanagari Character Recognition
摘要
This paper presents a novel method for creating a digital online handwritten Devanagari character dataset. Unlike existing datasets, which typically use scanned paper documents, this dataset was generated digitally through a Python-based canvas. The dataset includes 44 Devanagari characters, omitting 9 difficult-to-distinguish ones from the original 53. One hundred users contributed by drawing each character with a mouse, resulting in 44,000 images, each sized at 300 x 300 pixels. The images underwent resizing, noise reduction, and formatting into JPEG and CSV files, with a 70:30 training-to-testing split. A Convolutional Neural Network (CNN) was employed for character recognition using Keras’ Sequential model. The network included convolutional, pooling, ReLU activation, dropout, and dense layers, amounting to 2,387,116 trainable parameters. Each character was numerically labeled from 0 to 43, and 39,600 images were used for training, while 4,400 were tested. Results showed a recognition rate of 95.44% (4,199 correct out of 4,400) and a non-recognition rate of 4.58%. Performance metrics such as Precision, Recall, F1-score, Accuracy, and Specificity were analyzed using a confusion matrix. The model achieved accuracy rates of 95.00%, 95.45%, and 96.00% after 10, 20, and 30 epochs, respectively. After 30 epochs, the CNN reached an exceptional multi-class accuracy score of 0.99, setting a high benchmark for classification tasks. This innovative approach demonstrates the effectiveness of deep learning in recognizing Devanagari characters with high accuracy and precision.