A comprehensive review of statistical variants and enhancements of SMOTE oversampling method
摘要
The Synthetic Minority Oversampling Technique (SMOTE) has emerged as an effective and commonly used method for addressing class imbalance in machine learning modeling. This survey investigates the statistical variations of SMOTE, highlighting its enhancement types and techniques. Examining the statistical variants of SMOTE, specifically how synthetic instance production affects model performance, bias, overfitting, and generalization are the main ones. The research also emphasizes recent improvements, such as selection- and adjustment-based approaches. Overfitting, high-dimensional data, and generalization concerns are all covered. This survey provides an in-depth review of the usefulness and limitations of statistically based SMOTE and advises researchers and practitioners on selecting and customizing SMOTE techniques for specific datasets.