<p>In recent years, the proliferation of textual data sources has made it imperative to process large volumes of text quickly and accurately. Automated text summarisation has emerged as a viable solution to condense text documents while preserving crucial details. However, the unique characteristics of the Arabic language make Arabic text summarisation a particularly challenging task. This research focuses on Arabic text summarisation by incorporating Arabic Bidirectional Encoder Representations from Transformers (AraBERT), coupled with a score-based model and Part of Speech (POS) tagging features. Additionally, this study explores the impact of POS tagging by integrating it into several feature extraction and scoring steps. By combining AraBERT for pre-processing, TextRank for key phrase extraction, and POS tagging to enhance the scoring technique, we aim to effectively capture semantic parts of the text. Our proposed method was evaluated on the Essex Arabic Summary Corpus using the Recall-Oriented Understudy for Gisting Evaluation (ROUGE) metric, yielding encouraging results compared to existing approaches. Specifically, our approach achieved a precision rate of 59.4% with the EASC dataset. This high precision underscores the effectiveness of our method in retaining significant sentences and minimising redundancy. Furthermore, by comparing our approach's outcomes with other state-of-the-art methods, we demonstrated its effectiveness and potential to provide valuable insights for future enhancements. This study addresses the challenges inherent in Arabic text summarisation and sets the stage for future research in developing even more robust summarisation techniques.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Arabic text summarisation using Arabic bidirectional encoder representations from transformers-based models together with POS

  • Ghizlane Bourahouat,
  • Manar Abourezq,
  • Najima Daoudi

摘要

In recent years, the proliferation of textual data sources has made it imperative to process large volumes of text quickly and accurately. Automated text summarisation has emerged as a viable solution to condense text documents while preserving crucial details. However, the unique characteristics of the Arabic language make Arabic text summarisation a particularly challenging task. This research focuses on Arabic text summarisation by incorporating Arabic Bidirectional Encoder Representations from Transformers (AraBERT), coupled with a score-based model and Part of Speech (POS) tagging features. Additionally, this study explores the impact of POS tagging by integrating it into several feature extraction and scoring steps. By combining AraBERT for pre-processing, TextRank for key phrase extraction, and POS tagging to enhance the scoring technique, we aim to effectively capture semantic parts of the text. Our proposed method was evaluated on the Essex Arabic Summary Corpus using the Recall-Oriented Understudy for Gisting Evaluation (ROUGE) metric, yielding encouraging results compared to existing approaches. Specifically, our approach achieved a precision rate of 59.4% with the EASC dataset. This high precision underscores the effectiveness of our method in retaining significant sentences and minimising redundancy. Furthermore, by comparing our approach's outcomes with other state-of-the-art methods, we demonstrated its effectiveness and potential to provide valuable insights for future enhancements. This study addresses the challenges inherent in Arabic text summarisation and sets the stage for future research in developing even more robust summarisation techniques.