<p>Social media has provided the great opportunity for millions of internet users to express their opinions online. The online reviews have a huge potential to gain rich insight into an individual’s behavior towards an entity. Sentiment analysis (SA) is a widely accepted Natural Language Processing (NLP) technique that helps in analyzing these reviews. Feature extraction (FE) plays a key role in enhancing the performance of a sentiment classification model. Historically, the most popular techniques employed for FE have been Term Frequency-Inverse Document Frequency (TF-IDF), Word2Vec and GloVe. But these approaches are non-contextual and also domain-specific. The recent research studies have utilized state-of-the-art (SOTA) context-aware Bidirectional Encoder Representations from Transformers (BERT) embeddings for building SA models. However, BERT employs cross-encoder architecture in which it can produce word embeddings only. Sentence embeddings can be derived by averaging word embeddings. But this method is computationally expensive and does not yield optimal sentence embeddings (SEs). To address the shortcomings in the existing FE methods, this paper introduces a novel technique to extract domain-insensitive high-quality SEs directly via Sentence Transformer (ST) and proposes a general-purpose unified framework for SA. In the proposed framework, first we generate semantically rich SEs employing ST and then integrate these embeddings into three different machine learning algorithms including logistic regression, random forest and support vector machine. The proposed method is evaluated on balanced and imbalanced datasets across seven diverse domains on the basis of F1-score, precision, recall, accuracy and AUC values. It has outperformed the existing FE approaches and several SOTA studies. Also, the significant improvement in F1 scores, ranging from 4% to 7% above the SOTA BERT embeddings in case of imbalanced datasets, highlights the efficiency of the proposed method in dealing with imbalanced data.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Novel Transformer Based Unified Framework for Sentence-Level Sentiment Analysis

  • Khushboo Taneja,
  • Jyoti Vashishtha,
  • Saroj Ratnoo

摘要

Social media has provided the great opportunity for millions of internet users to express their opinions online. The online reviews have a huge potential to gain rich insight into an individual’s behavior towards an entity. Sentiment analysis (SA) is a widely accepted Natural Language Processing (NLP) technique that helps in analyzing these reviews. Feature extraction (FE) plays a key role in enhancing the performance of a sentiment classification model. Historically, the most popular techniques employed for FE have been Term Frequency-Inverse Document Frequency (TF-IDF), Word2Vec and GloVe. But these approaches are non-contextual and also domain-specific. The recent research studies have utilized state-of-the-art (SOTA) context-aware Bidirectional Encoder Representations from Transformers (BERT) embeddings for building SA models. However, BERT employs cross-encoder architecture in which it can produce word embeddings only. Sentence embeddings can be derived by averaging word embeddings. But this method is computationally expensive and does not yield optimal sentence embeddings (SEs). To address the shortcomings in the existing FE methods, this paper introduces a novel technique to extract domain-insensitive high-quality SEs directly via Sentence Transformer (ST) and proposes a general-purpose unified framework for SA. In the proposed framework, first we generate semantically rich SEs employing ST and then integrate these embeddings into three different machine learning algorithms including logistic regression, random forest and support vector machine. The proposed method is evaluated on balanced and imbalanced datasets across seven diverse domains on the basis of F1-score, precision, recall, accuracy and AUC values. It has outperformed the existing FE approaches and several SOTA studies. Also, the significant improvement in F1 scores, ranging from 4% to 7% above the SOTA BERT embeddings in case of imbalanced datasets, highlights the efficiency of the proposed method in dealing with imbalanced data.