Transformer-Based Automatic Sarcasm Detection System on Telugu Conversational Texts
摘要
Sarcasm is ofttimes used to convey an opposing view by utilizing positive or amplified positive words. As a result of this deliberate ambiguity, sarcasm detection is a crucial sentiment analysis challenge. This task becomes notably intricate when applied to a low-resource language such as Telugu, which requires more substantial linguistic resources and possesses complex morphological characteristics. The key challenge in this context revolves around acquiring a well-balanced and annotated dataset. In this chapter, we constructed a Telugu dataset consisting of 10,000 conversational texts, of which 5,000 are sarcastic, and the rest are not. Comprehensive annotation guidelines have been formed, and all sentences in the corpus have been meticulously annotated using multiple annotators. In addition, this chapter introduces an ensemble approach using three state-of-the-art transformer models (mBERT, DistilBERT, and MuRIL) to detect sarcasm in Telugu conversational sentences. The experimental findings clearly show the effectiveness of the proposed ensemble approach, which offers an improvement of approximately 1% over the baselines on the developed Telugu corpus.