Automatic Spelling Error Classification in Malayalam
摘要
Spelling errors, a commonplace phenomenon among students dominated by those with learning disabilities, call for effective systems that automatically and accurately identify and classify spelling mistakes. This research focused on developing an automatic spelling error classification tool to identify the Malayalam writing system’s phonological and orthographic spelling error categories. For the analysis of spelling errors in Malayalam, the study considered the input pair of real words and spelled words for extraction of six linguistic features: the difference in word length, edit distance, character overlapping percentage, the number of common phonemes, the number of common bigrams, and the number of common trigrams. After analysing spelling errors, machine learning techniques classify them. The research appraised the effectiveness of feature engineering apropos automatic spelling error classification in Malayalam. Results indicate that the fine-tuned Random Forest classifier achieved the highest classification accuracy of 71%, effectively distinguishing between the two error categories. The findings pinpoint the impact of linguistic features on accuracy and afford insights into the imperative of language technology tools for educational purposes.