<p>It is valuable to explore those hidden patterns from imbalanced data. In imbalanced data, skewed distribution of the classes makes minority classes to be hardly noticed. Existing classifiers easily suffer the perturbation caused by the skewed distribution so that they obtain unstable prediction and poor performance. Accordingly, to mitigate the perturbation, we utilize a score mechanism to employ a classifier being concerned about minority classes. Through calculating the conformity of the observed data, the score of the data conformity is obtained. And using the obtained score to classify highly imbalanced data. Following that, our classifier explores minority classes on the learning regions yielded by the score, instead of exploring them on the original data regions. Experimental results show that our classifier outperformed the competitors in classification performance and efficiency, moreover, it learned more compact boundaries separating minority classes from majority classes than the competitors did. Results also show that our classifier does not exhibit an exponential classification time at classifying large volume data with highly imbalanced ratio. We do not impose any restrictive classifier assumptions on both imbalance data and the calculation regarding the score of the data conformity. Additionally, we propose to sufficiently utilize the learning regions yielded by the score for better classification boundaries, instead of using the original data regions, since those hard-to-observe minority classes can be well perceived on the learning regions, meanwhile, there can tighten majority classes and minority classes so that the margins between them are significantly displayed.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Binary classification for imbalanced data using data conformity mechanism

  • Jian Zheng,
  • Shumiao Ren,
  • Jingyue Zhang,
  • Shiyan Wang,
  • Lin Li

摘要

It is valuable to explore those hidden patterns from imbalanced data. In imbalanced data, skewed distribution of the classes makes minority classes to be hardly noticed. Existing classifiers easily suffer the perturbation caused by the skewed distribution so that they obtain unstable prediction and poor performance. Accordingly, to mitigate the perturbation, we utilize a score mechanism to employ a classifier being concerned about minority classes. Through calculating the conformity of the observed data, the score of the data conformity is obtained. And using the obtained score to classify highly imbalanced data. Following that, our classifier explores minority classes on the learning regions yielded by the score, instead of exploring them on the original data regions. Experimental results show that our classifier outperformed the competitors in classification performance and efficiency, moreover, it learned more compact boundaries separating minority classes from majority classes than the competitors did. Results also show that our classifier does not exhibit an exponential classification time at classifying large volume data with highly imbalanced ratio. We do not impose any restrictive classifier assumptions on both imbalance data and the calculation regarding the score of the data conformity. Additionally, we propose to sufficiently utilize the learning regions yielded by the score for better classification boundaries, instead of using the original data regions, since those hard-to-observe minority classes can be well perceived on the learning regions, meanwhile, there can tighten majority classes and minority classes so that the margins between them are significantly displayed.