In the realm of image data classification, achieving balanced datasets is paramount for the efficacy of machine learning models. However, imbalanced datasets, wherein certain classes are underrepresented, pose a significant challenge in generating high-quality classifications. To address this issue, oversampling and undersampling techniques have emerged as popular strategies. We begin by elucidating the fundamental concepts behind oversampling and undersampling techniques, highlighting their respective advantages and limitations. Subsequently, we delve into an in-depth exploration of various oversampling methods, including random oversampling, Synthetic Minority Oversampling Technique (SMOTE) and generative adversarial network (GAN). Similarly, we scrutinize undersampling techniques such as random undersampling and Tomek links. Through extensive experimentation, we empirically evaluate the performance of both techniques on various degrees of imbalance. These findings are crucial as they underscore the practical implications of utilizing both techniques in image classification tasks. The experimental results demonstrate that oversampling methods tend to perform better in terms of low, medium, high, and very high imbalanced data problems. GAN has generally outperformed other sampling methods for various imbalance levels and SMOTE has shown good results in 20% imbalance induction. Random oversampling has maintained its performance accross various degrees of imbalance. However, for random oversampling and Tomek Links, the results continue to decline as imbalance level increases.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Comparative Analysis of Oversampling and Undersampling Techniques for Image Data Classification Across Varying Imbalance Levels

  • Deshant Singh,
  • Anurag Sharma

摘要

In the realm of image data classification, achieving balanced datasets is paramount for the efficacy of machine learning models. However, imbalanced datasets, wherein certain classes are underrepresented, pose a significant challenge in generating high-quality classifications. To address this issue, oversampling and undersampling techniques have emerged as popular strategies. We begin by elucidating the fundamental concepts behind oversampling and undersampling techniques, highlighting their respective advantages and limitations. Subsequently, we delve into an in-depth exploration of various oversampling methods, including random oversampling, Synthetic Minority Oversampling Technique (SMOTE) and generative adversarial network (GAN). Similarly, we scrutinize undersampling techniques such as random undersampling and Tomek links. Through extensive experimentation, we empirically evaluate the performance of both techniques on various degrees of imbalance. These findings are crucial as they underscore the practical implications of utilizing both techniques in image classification tasks. The experimental results demonstrate that oversampling methods tend to perform better in terms of low, medium, high, and very high imbalanced data problems. GAN has generally outperformed other sampling methods for various imbalance levels and SMOTE has shown good results in 20% imbalance induction. Random oversampling has maintained its performance accross various degrees of imbalance. However, for random oversampling and Tomek Links, the results continue to decline as imbalance level increases.