ANKA_INDI: A Comprehensive Stop Word Classification System for Indian Languages
摘要
Finding and getting rid of stop words is an important phase in text analysis and natural language processing. Eliminating stop words enhances search results in many Indian languages, such as Sanskrit, Bengali, Gujarati, and Marathi. Stop words are typically eliminated to increase part-of-speech tagging and named entity identification accuracy. The process of classifying texts is significantly, but unevenly, impacted by the elimination of stop words. Due to inconsistent standardization, complex morphology, lack of resources, and use dependence, stop word identification in Indian languages is difficult. Researchers have proposed many remedies to address these problems. Nevertheless, there is still a need for a system that can identify frequent Stop Words in Indian dialects. Various approaches to stop words are examined in this study, including hybrid, rule-based, corpus-based, and neural network-based ways. Along with emphasizing the limitations associated with Indian dialects, it offers a workable solution structure.