A new stemming model based on data generation to enhance Arabic information retrieval
摘要
Arabic information retrieval faces significant challenges due to the complex morphology of the Arabic text. Essential preprocessing steps, such as stemming and stop word removal, are crucial for an effective IR system. This research introduces a novel stemming model for Arabic text, leveraging Arabic morphological patterns. The model was trained on a systematically generated dataset of word-root pairs and evaluated on public benchmarks, demonstrating superior performance compared to existing techniques. Across all evaluation metrics, the proposed model achieved the highest effectiveness in improving Arabic information retrieval. The key contributions of this research include the creation of a large-scale dataset of word-root pairs derived from Arabic morphological patterns, the development and training of an advanced Arabic stemming model, and a comprehensive evaluation and comparative analysis of multiple stemming techniques in terms of stemming accuracy, and information retrieval.