错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Application of Mean-Variance Cloning Technique to Investigate the Comparative Performance Analysis of Classical Classifiers on Imbalance and Balanced Data

  • Friday Zinzendoff Okwonu,
  • Nor Aishah Ahad,
  • Joshua Sarduana Apanapudor,
  • Festus Irimisose Arunaye,
  • Olimjon Shukurovich Sharipov

摘要

Group imbalance data often occur in high dimensional practical classification problems where the number of attributes exceeds the number of instances. In this case, the researcher is faced with dual problems such as (i) biasedness towards the majority group over the minority group and (ii) dimensionality or singularity of the covariance matrix. In such situations, the classical classifiers dependent on the sample mean and covariance matrix are impracticable for classifications. This study focused on the effects of group imbalance data on \(n>p\) classification problems. First, we develop a procedure that could transform the minority group into the majority group before the classifiers are applied. This study aims to determine whether there are observable effects on the classifiers’ performance for imbalanced and balanced data sets. We also investigated whether the over and under-sampling influence the computational time of the classifiers. The results revealed that the Fisher linear classification method (FLCM) performed comparably for imbalanced and balanced data and outperformed the nearest mean classifier (NMC) and the independent classification rule (ICR). The study demonstrated that the MVCT effects on the classifiers are data-dependent. Therefore, the investigation showed that sample size balancing irrespective of the data dimension does not have a strong impact on the classifier’s performance. This analysis concludes that for the \(n>p\) classification problem, the FLCM classifier has comparable performance on the imbalanced and balanced data, similar results were observed for the NMC and ICR classifiers. The over and under-sampling of the data set has insignificant effects on the computational time of the classifiers.