An Empirical Study of Label Size Effect on Classification Model Accuracy Using a Derived Rule from the Holy Quran Verses
摘要
Machine Learning (ML) has become more and more significant in various applications, such as sentiment analysis and topic modelling, due to its ability to handle large volumes of text data and achieve high accuracy. Thus, the accuracy of the classification model using sentiment analysis has gained important heed recently because of its potential to provide valuable insights into customer preferences and public opinions. Accuracy is largely dependent on the quality and quantity of labelled data. This study aims to manifest the impact of label size on classification model accuracy by applying a derived rule from the Holy Quran verses which focuses on the useability of binary classification with two labels and comparing the accuracy models trained on a dataset labelled based on that rule and three-labelled dataset. The results show that the accuracy of the models trained on the binary-labelled dataset was higher than the accuracy of the models trained on the three-labelled dataset. The study’s findings will have implications for future research in ML models by applying the observed semantics from Quranic exegesis and analysis to improve the performance of ML models.