Sarcasm is ofttimes used to convey an opposing view by utilizing positive or amplified positive words. As a result of this deliberate ambiguity, sarcasm detection is a crucial sentiment analysis challenge. This task becomes notably intricate when applied to a low-resource language such as Telugu, which requires more substantial linguistic resources and possesses complex morphological characteristics. The key challenge in this context revolves around acquiring a well-balanced and annotated dataset. In this chapter, we constructed a Telugu dataset consisting of 10,000 conversational texts, of which 5,000 are sarcastic, and the rest are not. Comprehensive annotation guidelines have been formed, and all sentences in the corpus have been meticulously annotated using multiple annotators. In addition, this chapter introduces an ensemble approach using three state-of-the-art transformer models (mBERT, DistilBERT, and MuRIL) to detect sarcasm in Telugu conversational sentences. The experimental findings clearly show the effectiveness of the proposed ensemble approach, which offers an improvement of approximately 1% over the baselines on the developed Telugu corpus.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Transformer-Based Automatic Sarcasm Detection System on Telugu Conversational Texts

  • Ravi Teja Gedela,
  • Ujwala Baruah,
  • Badal Soni

摘要

Sarcasm is ofttimes used to convey an opposing view by utilizing positive or amplified positive words. As a result of this deliberate ambiguity, sarcasm detection is a crucial sentiment analysis challenge. This task becomes notably intricate when applied to a low-resource language such as Telugu, which requires more substantial linguistic resources and possesses complex morphological characteristics. The key challenge in this context revolves around acquiring a well-balanced and annotated dataset. In this chapter, we constructed a Telugu dataset consisting of 10,000 conversational texts, of which 5,000 are sarcastic, and the rest are not. Comprehensive annotation guidelines have been formed, and all sentences in the corpus have been meticulously annotated using multiple annotators. In addition, this chapter introduces an ensemble approach using three state-of-the-art transformer models (mBERT, DistilBERT, and MuRIL) to detect sarcasm in Telugu conversational sentences. The experimental findings clearly show the effectiveness of the proposed ensemble approach, which offers an improvement of approximately 1% over the baselines on the developed Telugu corpus.