AraXLM: New XLM-RoBERTa Based Method for Plagiarism Detection in Arabic Text
摘要
Significant advancements have been made over the previous decade in the development of plagiarism detection tools and software. This has expanded application beyond identification of plagiarized articles and research papers in academia, safeguarding authors’ literary rights across various domains, including digital marketing and libraries. However, the emergence of cross-language plagiarism poses both a significant challenge and a crucial area for further development. This primarily stems from the complexity of different languages, as well as the accuracy and interpretation of automatically translated words and sentences, which can prove more complex than the identification of plagiarism in a monolingual corpus. This position paper addresses the obstacles encountered in cross-language detection of plagiarism relating to Arabic-English and English-Arabic works. It outlines the impact of vowel marks on the semantic meaning of Arabic sentences. The paper also introduces a novel approach utilizing XLM-RoBERTa (XLMR) for enhancing the detection of plagiarism in Arabic text based on semantic similarity, referred to as AraXLM. This approach considers the influence of Arabic diacritical marks during text processing, as these can play a crucial role in improving translation accuracy and the interpretation of sentences. The proposed framework incorporates a cross-language approach utilizing features such as Facebook AI Similarity Search (FAISS), as well as semantic similarity and diacritization, to detect plagiarism in Arabic text at the sentence level of SemEval dataset. This research is ongoing, with the following stage being to evaluate the proposed framework by testing metrics such as Precision, Recall, and F-Measure. We anticipate that the employment of AraXLM for the detection of plagiarism in Arabic will address the challenges associated with, and enhance the effectiveness of, Natural Language Processing (NLP).