错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Oversampling Method Based Covariance Matrix Estimation in High-Dimensional Imbalanced Classification

  • Ireimis Leguen-de-Varona,
  • Julio Madera,
  • Hector Gonzalez,
  • Lise Tubex,
  • Tim Verdonck

摘要

Class imbalance is a common problem in (binary) classification problems. It appears in many application domains, such as text classification, fraud detection, churn prediction and medical diagnosis. A widely used approach to cope with this problem at the data level is the Synthetic Minority Oversampling Technique (SMOTE) which uses the K-Nearest Neighbors (KNN) algorithm to generate new, artificial instances in the minority class. It is however known that SMOTE is not ideal for high-dimensional data. Therefore, we propose an alternative oversampling strategy for imbalanced classification problems in high dimensions. Our approach is based on the sparse inverse covariance matrix estimated trough the Ledoit-Wolf method for high-dimensional data. The results show that our proposal has a competitive performance with respect to popular competitors.