错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Bi-SMOTE: a novel framework for handling imbalanced datasets using machine learning techniques

  • Onima Tigga,
  • Jaya Pal,
  • Debjani Mustafi

摘要

Working with Imbalance data in real-world problems is not so easy due to the different cardinality of classes. Several machine learning Techniques have been used to overcome this kind of problem for 100% original data for classification. In this research work, the proposed Bi-SMOTE approach on the imbalanced Red wine and Glass datasets using machine learning techniques, Random Forest (RF), K Nearest Neighbor (KNN), Logistic Regression (LR), and Decision Tree (DT) have been used. Moreover, preprocessing steps are done by Binarization and standardization. The accuracy and F1-score are calculated of without SMOTE / with SMOTE / proposed Bi-SMOTE approaches with classifiers RF, KNN, LR, and DT for 95% & 100% of original datasets. Here the Bi-SMOTE approach gives better results up to 12%, and up to 22% for accuracy and up to 29% and up to 28% for F1-score for the Red wine and the Glass dataset respectively. Moreover, mean squared error (MSE) is continuously decreased in these machine learning techniques with original data with SMOTE and Bi-SMOTE. Thus, the proposed approach Bi-SMOTE yields the highest values in all performance measures, including accuracy, and F1-score for 95% data than 100% data.