A Comparative Study of Machine Learning Algorithms on Datasets of Varying Sizes
摘要
This research investigates the nuanced performance dynamics of machine learning algorithms across datasets of varying sizes. Employing three distinct datasets characterized by different sample sizes, we implement four machine learning algorithms—Logistic Regression, K Nearest Neighbors (KNN), Naïve Bayes, and Random Forest. This paper delineates the application methodologies, hyperparameter tuning procedures, and the resultant performance outcomes. Our findings reveal that Logistic Regression and Naïve Bayes exhibit relatively minor sensitivity to dataset size fluctuations. In stark contrast, KNN is markedly influenced by dataset size, showcasing significant performance variations. Notably, Random Forest demonstrates consistently superior performance across the four machine learning models considered, particularly excelling with the largest datasets. The implications of these observations are discussed, contributing to a comprehensive understanding of the interplay between machine learning algorithms and dataset characteristics. The effectiveness of a machine learning algorithm does not rely solely on expanding the dataset size. Our research indicates that once a certain threshold is reached, merely increasing the dataset size does not proportionally improve performance.