NLP-Based Processing of Gujarati Compound Word Sandhi’s Generation and Segmentation
摘要
In this research paper, we focused on the contribution of rule-based “sandhi” splitting and joining methods for the Gujarati language, a prominent Indian language with a complex linguistic structure. “Sandhi” represents a crucial aspect of the language grammar, where words undergo changes when combined, affecting both written and spoken communication. In this study, we explore traditional rule-based techniques for splitting and joining Gujarati compound words, an integral part of the language’s morphology. We examine the effectiveness of these methods in enhancing text processing and comprehension, particularly in the perspective of NLP applications. In this study, we delve into the linguistic intricacies of Gujarati sandhi and develop rule-based algorithms to effectively split and join words according to established grammatical rules. We explore the application of these methods in various natural language processing tasks, such as text analysis, machine translation, and information retrieval. Sandhi Splitting (Vicched) is a trickier task then sandhi joining for given its ordinariness and context dependency. Furthermore, we compare the performance of rule-based sandhi methods with other existing approaches, including machine learning-based techniques, to evaluate their effectiveness and applicability in different contexts. Additionally, we discuss potential enhancements and future directions for incorporating these rule-based methods into modern computational linguistics and language processing frameworks.