A Hybrid System for Robust Santali-English Language Identification
摘要
This study investigates whether Santali words can be differentiated from English based on linguistic patterns by using a Santali dataset for training instead of explicit English data. Character n-grams with SVM, Random Forest, Naive Bayes, CNN, LSTM, various ensemble models, and a hybrid model are among the classical, ensemble, and deep learning techniques that we use. The findings demonstrate how well Santali characteristics are captured by character n-grams, allowing for generalization to English words. Although random forest and ensemble models are good substitutes for environments with limited resources, CNN attains better accuracy. The hybrid model combines several classifiers to further improve performance. The error analysis shows that while longer words with different properties are easier to identify, small Santali terms that overlap with English present categorization issues. This research contributes to language identification for under-resourced languages, providing a foundation for Santali NLP tools and highlighting the potential of deep learning and hybrid models in this domain.