From Dataset Creation to Recognition Precision: A Comprehensive Study on Machine Learning Algorithms for Tulu Script
摘要
A notable gap exists in the market for effective OCR solutions for Indian languages, especially Tulu, which is widely spoken in South Karnataka. This is due to the fact that the majority of current research in optical character recognition (OCR) systems focuses on languages other than Indian. This paper presents a specific OCR system that can recognize vowels and consonants, among other fundamental characters, in photos written in Tulu script. The system’s adaptability allows it to accommodate a large variety of font sizes and styles. This work is significant because it addresses the deficiency of OCR programs specifically designed for Indian languages where there is lack of standardized datasets. The proposed work on From Dataset Creation to Recognition Precision Study on Machine Learning Algorithms for Tulu Script System’s design, workings, and potential applications are all thoroughly discussed. It is important to notice that a classifier-level fusion technique is included since it makes clear how it might improve recognition accuracy, particularly in light of the nuanced background of Tulu individuals. The proposed method’s objective is to improve OCR technology for Indian languages by providing a solution that is especially made to satisfy Tulu script’s requirements.