<p>Arabic research on climate sentiment analysis remains limited due to linguistic diversity, limited resources, and the scarcity of annotated datasets. Existing approaches primarily rely on questionnaires or general-purpose sentiment analysis models, which often fail to capture domain-specific sentiment expressions. To address this gap, this paper introduces <i>ClimaSentAR</i>, the first large-scale Arabic climate-focused sentiment dataset comprising 23,957 samples from the X platform. Native Arabic speakers annotate the dataset with a moderate inter-annotator agreement of 0.45. It contains 5510 positive, 8140 negative, and 10,307 neutral instances. To validate the importance and robustness of <i>ClimaSentAR</i> and address the gap in Arabic climate change sentiment analysis, this study fine-tunes Arabic-centric pre-trained language models and large language models. In contrast to prior work based solely on general-purpose datasets, the proposed approach systematically examines both in-domain performance and cross-domain generalization, while incorporating adaptation strategies including Domain-Adaptive Pre-training, AdapterFusion, and data augmentation. Experimental results demonstrate that climate-specific fine-tuning substantially improves in-domain climate sentiment performance, with AraBERT achieving the highest macro F1 score of 0.753 on climate-related data. Cross-domain evaluations resulted in degraded model performance, with climate-tuned models achieving a macro F1 score of up to 0.599, compared to less than 0.48 for general-purpose models. To support reproducibility and future research, the fine-tuned AraBERT model is publicly available on Hugging Face with an interactive web interface for testing and further fine-tuning by researchers and practitioners.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The ClimaSentAR dataset for evaluating cross-domain transferability of large language models in climate sentiment analysis

  • Sanaa Kaddoura,
  • Mariam Alzaabi,
  • Reem Nassar

摘要

Arabic research on climate sentiment analysis remains limited due to linguistic diversity, limited resources, and the scarcity of annotated datasets. Existing approaches primarily rely on questionnaires or general-purpose sentiment analysis models, which often fail to capture domain-specific sentiment expressions. To address this gap, this paper introduces ClimaSentAR, the first large-scale Arabic climate-focused sentiment dataset comprising 23,957 samples from the X platform. Native Arabic speakers annotate the dataset with a moderate inter-annotator agreement of 0.45. It contains 5510 positive, 8140 negative, and 10,307 neutral instances. To validate the importance and robustness of ClimaSentAR and address the gap in Arabic climate change sentiment analysis, this study fine-tunes Arabic-centric pre-trained language models and large language models. In contrast to prior work based solely on general-purpose datasets, the proposed approach systematically examines both in-domain performance and cross-domain generalization, while incorporating adaptation strategies including Domain-Adaptive Pre-training, AdapterFusion, and data augmentation. Experimental results demonstrate that climate-specific fine-tuning substantially improves in-domain climate sentiment performance, with AraBERT achieving the highest macro F1 score of 0.753 on climate-related data. Cross-domain evaluations resulted in degraded model performance, with climate-tuned models achieving a macro F1 score of up to 0.599, compared to less than 0.48 for general-purpose models. To support reproducibility and future research, the fine-tuned AraBERT model is publicly available on Hugging Face with an interactive web interface for testing and further fine-tuning by researchers and practitioners.