A consistency analysis on four evaluation metrics for classifying imbalanced data
摘要
Several metrics have been proposed to replace accuracy for evaluating the classification performance on imbalanced data. The properties of the metrics are different, and which one should be adopted for evaluation or whether their evaluation results are consistent is still inconclusive. Four popular metrics AUC, F-measure, G-mean, and MCC are considered in this study to investigate their evaluation consistency. Simulation studies are first executed to explore the impact of imbalanced ratio on the properties of F-measure, G-mean, and MCC. The simulation results show that when the imbalanced ratios for all data sets are not less than eight, the three threshold metrics have very strong linear relationship and very high consistency rate. The lower and upper limits of AUC for any given confusion matrix are derived to explore its properties by regression analysis. The experimental results on 40 real data sets confirm the findings for the three threshold metrics, and the inconsistency rate between AUC and each of the three threshold metrics is significantly larger than 25%. The three threshold metrics generally result in the same conclusion that is usually different from the one obtained from AUC.