Predicting agricultural production is a very important challenge given global demography. However, climate change is one of the factors if not the main key that often skews this agricultural production prediction. Finding a machine learning method to better predict rice production in the Niger River Inner Delta represents an interesting challenge. We use a rice production (Climate_Rice) dataset in this paper. The Climate_Rice dataset is found to be a dataset with a class imbalance problem. In this paper thirty others imbalanced datasets were treated to study the class imbalance phenomenon in a dataset. We also propose a data-level method of solving the class imbalance problem called: Synthetic Minority Oversampling Technique in Stages (SMOTEiS). Our method has been used for Multilayer Perceptron (MLP), Logistic Regression (LR) and Support Vector Machine (SVM) classifiers. We compared our method with SMOTE, AdaboostM1, Bagging, CSForest and WiSARD to study its competitiveness. Accuracy being not a good metric for the class imbalance problem, we use three other metrics in this paper such as F-measure, G-mean and MCC (Matthew’s Correlation Coefficient). The results of these metrics show that our method is largely competitive with well-known methods for the class imbalance problem in a dataset.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Imbalanced Data Classification Using Synthetic Minority Oversampling Technique in Stages for a Rice Dataset

  • Moussa Diallo,
  • Abdoulaye Sidibé,
  • Djibril Diarra

摘要

Predicting agricultural production is a very important challenge given global demography. However, climate change is one of the factors if not the main key that often skews this agricultural production prediction. Finding a machine learning method to better predict rice production in the Niger River Inner Delta represents an interesting challenge. We use a rice production (Climate_Rice) dataset in this paper. The Climate_Rice dataset is found to be a dataset with a class imbalance problem. In this paper thirty others imbalanced datasets were treated to study the class imbalance phenomenon in a dataset. We also propose a data-level method of solving the class imbalance problem called: Synthetic Minority Oversampling Technique in Stages (SMOTEiS). Our method has been used for Multilayer Perceptron (MLP), Logistic Regression (LR) and Support Vector Machine (SVM) classifiers. We compared our method with SMOTE, AdaboostM1, Bagging, CSForest and WiSARD to study its competitiveness. Accuracy being not a good metric for the class imbalance problem, we use three other metrics in this paper such as F-measure, G-mean and MCC (Matthew’s Correlation Coefficient). The results of these metrics show that our method is largely competitive with well-known methods for the class imbalance problem in a dataset.