As the world is growing, the amount of data generated has been growing proportionately. Classification of the large amount of data generated presents multiple issues including skewness and sparsity of the available data which may result in a significant imbalance in the classes. Several techniques such as sampling, cost-sensitive learning, and the use of ensemble methods have been researched to tackle this issue of class imbalance or distribution. Cost-sensitive learning is a potent approach in terms of dealing with imbalanced data that may be used to enhance the performance of the machine learning model, lower the risk of false negatives, and improve the interpretability of the model. This study presents an empirical approach by using cost-sensitive learning on various predictive models such as binary logistic regression, decision tree, and support vector machine classifiers. An investigation is carried out on the differences in the accuracy of the classification when each of these machine learning models is trained with and without assigning misclassification costs to the classes of the dataset. The experiment conducted to validate the research presents conflicting results as the accuracy of the classification models decreases by providing a cost, but a significant increase in recall is observed upon adding the misclassification costs to the model.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Training Cost-Sensitive Classifiers to Tackle Imbalanced Data

  • Utkarsh Kejriwal,
  • Komal Arora

摘要

As the world is growing, the amount of data generated has been growing proportionately. Classification of the large amount of data generated presents multiple issues including skewness and sparsity of the available data which may result in a significant imbalance in the classes. Several techniques such as sampling, cost-sensitive learning, and the use of ensemble methods have been researched to tackle this issue of class imbalance or distribution. Cost-sensitive learning is a potent approach in terms of dealing with imbalanced data that may be used to enhance the performance of the machine learning model, lower the risk of false negatives, and improve the interpretability of the model. This study presents an empirical approach by using cost-sensitive learning on various predictive models such as binary logistic regression, decision tree, and support vector machine classifiers. An investigation is carried out on the differences in the accuracy of the classification when each of these machine learning models is trained with and without assigning misclassification costs to the classes of the dataset. The experiment conducted to validate the research presents conflicting results as the accuracy of the classification models decreases by providing a cost, but a significant increase in recall is observed upon adding the misclassification costs to the model.