<p>Deep learning has significantly advanced the question-answering (QA) systems across various sectors. However, Arabic-language systems for Hajj-related fatwas (non-binding Islamic legal opinions issued by muftis) remain underdeveloped. This paper introduces Hajj-FQA, a benchmark Arabic dataset specifically designed to develop HajjBot - a specialized chatbot for fatwas QA during the annual Hajj pilgrimage. The dataset captures the unique linguistic and jurisprudential characteristics of pilgrims’ inquiries, enabling accurate, domain-specific responses. We present a comprehensive quantitative analysis of the dataset’s construction methodology and its distinctive question-answer patterns. Evaluation using multilingual and Arabic-specific language models across three tasks - machine reading comprehension (MRC), duplicate question detection (DQD), and duplicate answer detection (DAD) - with 10-fold cross-validation demonstrates the practical utility of Hajj-FQA. Results show exceptional performance in classification tasks (AraBERTv0.2 achieved a precision score of 99.19% for DQD and 99.26% for DAD) and strong extractive answering capability with an <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="44443_2025_128_Article_IEq1.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="56" /> </InlineMediaObject> <EquationSource Format="TEX">\(F\text {-score}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>F</mi> <mtext>-score</mtext> </mrow> </math></EquationSource> </InlineEquation> of 72.78%. While generative performance reached <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="44443_2025_128_Article_IEq2.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="66" /> </InlineMediaObject> <EquationSource Format="TEX">\(\text {BERT-}F\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mtext>BERT-</mtext> <mi>F</mi> </mrow> </math></EquationSource> </InlineEquation> score of 71.4% (AraBART), MRC variability highlights challenges in religious reasoning. These findings establish Hajj-FQA as both: (1) a critical resource for developing specialized fatwa chatbots like HajjBot, and (2) a benchmark for Arabic religious QA systems. The dataset directly addresses the urgent need for accurate, automated fatwa assistance during Hajj, while providing insights for future improvements in Islamic NLP applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hajj-FQA: A benchmark Arabic dataset for developing question-answering systems on Hajj fatwas

  • Hayfa A. Aleid,
  • Aqil M. Azmi

摘要

Deep learning has significantly advanced the question-answering (QA) systems across various sectors. However, Arabic-language systems for Hajj-related fatwas (non-binding Islamic legal opinions issued by muftis) remain underdeveloped. This paper introduces Hajj-FQA, a benchmark Arabic dataset specifically designed to develop HajjBot - a specialized chatbot for fatwas QA during the annual Hajj pilgrimage. The dataset captures the unique linguistic and jurisprudential characteristics of pilgrims’ inquiries, enabling accurate, domain-specific responses. We present a comprehensive quantitative analysis of the dataset’s construction methodology and its distinctive question-answer patterns. Evaluation using multilingual and Arabic-specific language models across three tasks - machine reading comprehension (MRC), duplicate question detection (DQD), and duplicate answer detection (DAD) - with 10-fold cross-validation demonstrates the practical utility of Hajj-FQA. Results show exceptional performance in classification tasks (AraBERTv0.2 achieved a precision score of 99.19% for DQD and 99.26% for DAD) and strong extractive answering capability with an \(F\text {-score}\) F -score of 72.78%. While generative performance reached \(\text {BERT-}F\) BERT- F score of 71.4% (AraBART), MRC variability highlights challenges in religious reasoning. These findings establish Hajj-FQA as both: (1) a critical resource for developing specialized fatwa chatbots like HajjBot, and (2) a benchmark for Arabic religious QA systems. The dataset directly addresses the urgent need for accurate, automated fatwa assistance during Hajj, while providing insights for future improvements in Islamic NLP applications.