错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Investigating the Effects of Applying Different Text Pre-processing on the Performance of Sentiment Analysis for Malay Document Corpus

  • Rayner Alfred,
  • Elly Mazlin Binti Rahim,
  • Rayner Henry Pailus

摘要

In this era where social media networking and internet applications keep growing, people can easily express their opinions and feelings towards anything, either on products or services, in social media. These online reviews are the primary source in text analysis tasks, such as a sentiment analysis which is used to understand people’s sentiments towards something. Usually, online reviews are raw data that come in unstructured form and have a lot of noise. It needs to be processed and cleaned so that it can be used for further analysis. So far, there are limited works conducted on investigating the effects of applying text preprocessing tasks in the sentiment analysis for Malay language corpus, despite having enough data to work on it. Thus, this study investigates the impacts of text preprocessing to sentiment analysis on the Malay language using different approaches. The focus is more on handling the emojis and emoticons, typographical errors, repeated characters in the word, and the stemming process. The dataset, which is related to product reviews, was collected manually using a web scraping technique. The Support Vector Machine (SVC) classifier was applied to assess the effects of different text preprocessing approaches on the quality of sentiment classification based on accuracy, precision, recall, and F1-score. The results showed that the text preprocessing tasks have significant impacts on the performance accuracy of the Malay sentiment analysis.