错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Toward an efficient extractive Arabic text summarisation system based on Arabic large language models

  • Ghizlane Bourahouat,
  • Manar Abourezq,
  • Najima Daoudi

摘要

Automated text summarisation is challenging for the Arabic language due to the performance limitations of existing systems. Hence, our paper focuses on extractive Arabic text summarisation by incorporating Arabic large language models (LLMs) such as AraT5, AraGPT2, and AraBART compared to unsupervised techniques, namely TextRank and KeyBERT. Moreover, this work aims to investigate the impact as well as the performance of Arabic LLMs by testing several models. Our methodology is assessed on the Essex Arabic Summary Corpus (EASC), containing 153 Arabic text article and their corresponding 765 human-generated extractive summaries. We employ the ROUGE metric and manual human evaluation to assess the outcomes, and our approach demonstrates promising results compared to alternative existing methods. The research provides a comparative analysis about the performance of LLMs and traditional models. Through program-based evaluation, our approach achieves noteworthy precision levels. Specifically, when utilising the AraT5 model, we achieve a precision of 53.5%, while the AraGPT2 model achieves a precision of 53%. Additionally, our human-based evaluation, which captures users' perspectives, reveals that our proposed approach can generate coherent Arabic summaries that are easy to read and effectively capture the main idea of the input text.