Performance Based Comparative Analysis of Naïve Bayes Variants for Text Classification
摘要
With the high consumption of digital data, the problem of unorganized textual data created major challenges in today’s scenario. To overcome this issue of unorganized textual information, we perform document classification which is divided into four major phases, i.e., pre-processing, feature selection, model training, and model testing. In this paper, we have selected four types of feature vectors, three with weighting techniques and one without weighting. Performance has been tested on these four feature vectors with five variances of Naïve Bayes classifiers out of which two were not able to perform training due to the sparseness in the dataset and the rest three performed well. In the reported result of the experiment, F1-macro score and the accuracy of Complement Naïve Bayes using term frequency weighting scheme is 0.807901362 and 0.821163038 which outperform all the other feature sets with all the variance of Naïve Bayes classifiers. In terms of time consumption for training and testing again, the performance of Complement Naïve Bayes using term frequency weighting scheme found best.