Sentiment Analysis usually classifies a sentence by label it with positive or negative. Previous to apply a classification model, we need a big corpus with many diverse cases. In this paper we run different paraphrasing techniques over the IMDB reviews dataset increasing the number of sentences. We tested such augmented corpus using a BERT classification model. The results: without paraphrasing, the BERT model obtained 90% of f-measure; and with paraphrasing we obtained between 99.56% and 99.86%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing BERT Classification by Increasing Training Corpus Using Paraphrasing

  • Carlos E. Atencio-Torres,
  • Julio A. Vera-Sancho,
  • Pablo C. Calcina-Ccori,
  • Alvaro H. Mamani-Aliaga

摘要

Sentiment Analysis usually classifies a sentence by label it with positive or negative. Previous to apply a classification model, we need a big corpus with many diverse cases. In this paper we run different paraphrasing techniques over the IMDB reviews dataset increasing the number of sentences. We tested such augmented corpus using a BERT classification model. The results: without paraphrasing, the BERT model obtained 90% of f-measure; and with paraphrasing we obtained between 99.56% and 99.86%.