错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dictionary-based extraction of hyperbole and swear words for sarcasm detection in Indonesian Tweets

  • Novitasari Arlim,
  • Al Hafiz Akbar Maulana Siagian,
  • Slamet Riyanto,
  • Rodiah Rodiah,
  • Siti Kania Kushadiani,
  • Shidiq Al Hakim,
  • Retno Asihanti Setiorini,
  • Niken Fitria Apriani,
  • Rini Arianty,
  • Diana Tri Susetianingtias

摘要

Detecting sarcasm in texts presents a formidable challenge due to the absence of clear signals, unlike in verbal conversations. Relying on hashtags for sarcasm detection may lack accuracy and standardization, while human annotation of sarcasm can be subjective. Moreover, detecting sarcasm in a low-resource language like Indonesian presents unique challenges. In this work, we utilize hyperbole, i.e., interjection (INJ), intensifier (INS), capital letters (CL), elongated words (EW), and punctuation marks (PUNC), as features for classifiers to detect sarcasm in Indonesian texts. We also propose using swear words (SW) as features to deal with this task. As our other proposal in this work, we consider a dictionary-based method to extract these features more effectively. To evaluate our work, we use a dataset containing Indonesian tweets collected from Drone Emprit. Experimental results show that incorporating hyperbolic features with SW improves sarcasm detection performance compared to using these features separately. Our proposed feature extraction method also produces better classification results than the baseline.