Sentiment analysis in product reviews in Thai language
摘要
The volume of customer reviews on e-commerce platforms is experiencing a rapid surge. Low-resource languages like Thai face challenges in sentiment polarity classification. With the emergence of transformer models pre-trained in diverse languages, new possibilities have emerged. We employed XLM-Roberta, BERT, and two pre-trained weights of WangchanBERTa. In the experimental results, after finetuning with our dataset, WangchanBERTa achieved a maximum accuracy of 85%. Furthermore, this model was considered the best, with performance metrics ranging from 66 to 93% for precision, recall, and the highest F1-score on the test Thai reviews dataset. Additionally, by observing the characteristics of the Thai text reviews before employing transformer-based models, we found that data cleansing with PyThaiNLP contributed to a modest 2% improvement in the performance of the WangchanBERTa model.