错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Rough Sets in Imbalanced Data Problem: An Improving Oversampling Process

  • Sara A. Shehab,
  • Ashraf Darwish,
  • Aboul Ella Hassanien

摘要

The problem of imbalanced data continues to be one of the most significant study topics. The most recent studies and in-depth studies showed that the intrinsic complexity of data, rather than just the underrepresented classes, is what is mostly causing performance loss in machine learning processes. The issue of learning from unbalanced data has become one that is significant and difficult. The performance of the classifier is significantly diminished because of the complicated distribution of the data, particularly minor disjuncts, noise, and class overlaps. Consequently, a variety of options were put forth. This work provides a data-level and algorithm-level approach that combines rough sets with SMOTE. The approach was used to solve imbalanced issues, and the datasets’ nominal properties were used to describe them. The proposed work is compared to various preprocessing techniques and rated it. Six different unbalanced datasets are used in the evaluation and three different machine learning algorithms. In the breast cancer dataset, the accuracy improved from 88 to 95%, 64.29 to 92.65%, and 95.71 to 97.06% with Gaussian-NB, Logistic Regression, and Decision tree, respectively. The combination of SMOTE and Rough Set Theory for preprocessing within the context of imbalanced datasets produced good average outcomes, according to the findings of our experimental investigation.