The burgeoning growth of data in recent years has led to datasets becoming increasingly larger, which poses challenges for the development of machine learning models. The complexity of contemporary machine learning algorithms often results in slower training times and longer development periods. To address these issues, this paper proposes a novel numerosity reduction algorithm that leverages the Central Limit Theorem (CLT) to reduce the number of data points for training machine learning models while preserving interpretability and conforming to a Gaussian distribution. The proposed numerosity reduction algorithm is expected to enhance the training time of binary classification models and improve the performance of algorithms that rely on the assumption that the training features are Gaussian distributed. Consequently, this technique may facilitate the development of large and complex machine learning algorithms, making them more feasible to train and deploy. The aim is to decrease the number of data points (cloud points) used for training machine learning models, while preserving interpretability and transforming the distribution of the variables to Gaussian. This is particularly advantageous as several algorithms assume data to be normally distributed. By harnessing the properties of CLT, it is expected that training times for binary classification algorithms will be enhanced, as well as the performance of algorithms that rely on gradient descent for optimization. This speed in training is pivotal for the development of large and complex machine learning algorithms.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Analysis of the Central Limit Theorem-Driven Numerosity Reduction Technique: Implications for Binary Classification Algorithm Performance and Training Duration

  • Afizan Azman,
  • Husna Sarirah Husin,
  • Norhidayah Hamzah,
  • Swee King Phang,
  • Sumendra Yogarayan,
  • R. Lalitha

摘要

The burgeoning growth of data in recent years has led to datasets becoming increasingly larger, which poses challenges for the development of machine learning models. The complexity of contemporary machine learning algorithms often results in slower training times and longer development periods. To address these issues, this paper proposes a novel numerosity reduction algorithm that leverages the Central Limit Theorem (CLT) to reduce the number of data points for training machine learning models while preserving interpretability and conforming to a Gaussian distribution. The proposed numerosity reduction algorithm is expected to enhance the training time of binary classification models and improve the performance of algorithms that rely on the assumption that the training features are Gaussian distributed. Consequently, this technique may facilitate the development of large and complex machine learning algorithms, making them more feasible to train and deploy. The aim is to decrease the number of data points (cloud points) used for training machine learning models, while preserving interpretability and transforming the distribution of the variables to Gaussian. This is particularly advantageous as several algorithms assume data to be normally distributed. By harnessing the properties of CLT, it is expected that training times for binary classification algorithms will be enhanced, as well as the performance of algorithms that rely on gradient descent for optimization. This speed in training is pivotal for the development of large and complex machine learning algorithms.