Enhancing Book Recommendations on GoodReads: A Data Mining Approach Based Random Forest Classification
摘要
With the rise of technology, new ways of finding books have emerged beyond traditional bookstores. Websites like www.goodreads.com allow readers to share their book reviews and ratings. This study uses the data from GoodReads to find the best way to suggest books to readers. Employing data mining classification, four methods - Random Forest, Naive Bayes, K-Nearest Neighbor, and Support Vector Classifier - were examined. Performance evaluation was conducted using accuracy, F-measure, recall, and precision metrics derived from the confusion matrix. Interestingly, the Random Forest algorithm stood out with remarkable results. It achieved 99.91% accuracy, 100% precision, 92% recall, a 95% F1-score, and a slight 0.09 average error. These impressive outcomes highlight the algorithm’s effectiveness in predicting user preferences and offering personalized book recommendations. Additionally, the study compared the Random Forest approach with the baseline methods, showing its clear superiority. This research showcases the promising potential of Random Forest in improving the GoodReads book recommendation system. Using the random forest classifier proved effective in predicting user preferences and generating relevant book recommendations, offering a promising approach to enhance personalized reading experiences, fitting well with changing reading habits in the digital era.