Automatic Identification of Meimayakkam in Tamil Words Using Rule Based and Transfer Learning Approaches
摘要
Over 50 traditional Tamil grammar books are in existence, with Tholkappiyam standing out as the most significant. Written in 14 BC, it comprises three chapters: Ezhuththathikaaram ( ), Chollathikaaram ( ), and Porulathikaaram ( ). These chapters delve into the structures of the Tamil language, encompassing the concept of Meimayakkam , which elucidates the correct spelling usage of Tamil words—an imperative facet for grasping the intricacies of the language. Tholkappiyar expounds on this concept in two sub-chapters, Nuunmarapu ( ) and Mozhimarapu ( ), which encompass 82 rules, including 12 specifically dedicated to Meimayakkam. This paper introduces two methods for automatically detecting spelling mistakes, utilizing the Meimayakkam Rules Multilingual BERT Model, a Large Language Model. These approaches are then compared in the context of the spell-checking task. Additionally, a mobile application has been developed based on the Meimayakkam rules, serving as an educational tool for students learning Tamil. This study is groundbreaking as it marks the first examination of Tholkaappiyar rules from a linguistic perspective, upholding the purity of Tamil words and distinguishing them from vernacular words.