In-Depth Analysis and Challenges of Handwritten Telugu Character Recognition with an Improved CNN Model for Solutions
摘要
The Indian Constitution recognizes Telugu, Tamil, Malayalam, and Kannada as significant languages. Approximately 90 million people worldwide speak Telugu, a language from South India. Telugu OCR, or optical character recognition, has several uses, such as enabling digital books and unstructured documents, which in turn improves communication between people. In contrast to regional languages like Telugu, Tamil, Malayalam, etc., OCR systems are extremely well trained for international languages like English and German. Telugu has a large number of different characters, which presents a significant obstacle to the development of OCR. In this study, we address this in two ways: (i) a database containing handwritten Telugu characters and (ii) an enhanced CNN model designed to identify scanned Telugu characters.