<p>Oversampling algorithms improve imbalanced dataset classification by increasing the quantity of the minority class to reduce the domination of the majority class. The existing oversampling methods only focus on solving the imbalance problem between classes, often ignoring the imbalance problem within the class, and even producing noise or unnecessary samples. The weighted minority cluster oversampling algorithm (WMCOA) was proposed to solve this problem. This method can avoid the generation of new noise samples, and effectively solve the problem of intra-class and inter-class imbalance, so as to make the data set samples achieve a better-balanced proportion. In the experiments of 20 imbalanced data sets, support vector machine, K-nearest neighbor, and random forest were used as classifiers. The experimental results show that compared with several commonly used imbalanced data processing algorithms, the proposed algorithm can effectively improve the performance of imbalanced data. The WMCOA algorithm has achieved better results in <i>F</i>-measure, <i>G</i>-mean, and AUC evaluation indicators.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An oversampling method based on the weighting of minority class clusters

  • Yunbin He,
  • Chenglong Li,
  • Fuwei Guo

摘要

Oversampling algorithms improve imbalanced dataset classification by increasing the quantity of the minority class to reduce the domination of the majority class. The existing oversampling methods only focus on solving the imbalance problem between classes, often ignoring the imbalance problem within the class, and even producing noise or unnecessary samples. The weighted minority cluster oversampling algorithm (WMCOA) was proposed to solve this problem. This method can avoid the generation of new noise samples, and effectively solve the problem of intra-class and inter-class imbalance, so as to make the data set samples achieve a better-balanced proportion. In the experiments of 20 imbalanced data sets, support vector machine, K-nearest neighbor, and random forest were used as classifiers. The experimental results show that compared with several commonly used imbalanced data processing algorithms, the proposed algorithm can effectively improve the performance of imbalanced data. The WMCOA algorithm has achieved better results in F-measure, G-mean, and AUC evaluation indicators.