<p>The advancement of NLP has made significant strides in sentiment style transfer, modifying the linguistic style of a text while preserving its content. However, most existing datasets are non-parallel and focus on English, neglecting low-resource languages like Arabic. The lack of comprehensive Arabic parallel datasets has hindered the development and evaluation of robust sentiment transfer models for Arabic. To address this, we introduce MA’AKS, a novel Arabic parallel dataset for sentiment style transfer. MA’AKS consists of 5k sentences in modern standard Arabic with positive/negative sentiments. Each sentence is meticulously annotated to ensure high-quality parallel sentiment pairs, supporting both supervised and unsupervised learning. To benchmark the dataset, we evaluated AceGPT, JAIS, and Llama-3 LLMs on Arabic sentiment transfer with different learning settings, including zero-shot, few-shot, and fine-tuning. By publicly releasing MA’AKS, annotation guidelines, and experiment code, we aim to advance research on Arabic sentiment transfer and contribute to the NLP community.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Ma’aks: manually-curated parallel dataset for Arabic text sentiment swap

  • Raed Mughaus,
  • Shadi Abudalfa,
  • Hamzah Luqman,
  • Fahad Abdu,
  • Mohammed AlAli,
  • Nawaf Al-Dowayan,
  • Ahmed Abdelali

摘要

The advancement of NLP has made significant strides in sentiment style transfer, modifying the linguistic style of a text while preserving its content. However, most existing datasets are non-parallel and focus on English, neglecting low-resource languages like Arabic. The lack of comprehensive Arabic parallel datasets has hindered the development and evaluation of robust sentiment transfer models for Arabic. To address this, we introduce MA’AKS, a novel Arabic parallel dataset for sentiment style transfer. MA’AKS consists of 5k sentences in modern standard Arabic with positive/negative sentiments. Each sentence is meticulously annotated to ensure high-quality parallel sentiment pairs, supporting both supervised and unsupervised learning. To benchmark the dataset, we evaluated AceGPT, JAIS, and Llama-3 LLMs on Arabic sentiment transfer with different learning settings, including zero-shot, few-shot, and fine-tuning. By publicly releasing MA’AKS, annotation guidelines, and experiment code, we aim to advance research on Arabic sentiment transfer and contribute to the NLP community.