Imbalanced datasets present a significant challenge in supervised learning, particularly in image classification tasks where the number of instances varies widely among different classes, leading to reduced accuracy and bias toward more prevalent classes. This paper addresses the challenge posed by imbalanced datasets, exemplified by Caltech101, which includes a diverse range of instances per class, from as few as 31 to as many as 800. We introduce a robust technique combining artificial data augmentation with a pre-trained VGG16 architecture to significantly enhance classification accuracy. Our approach employs image augmentation techniques, including rotation, flipping, and noise injection, as a preprocessing step to enrich the dataset and mitigate class imbalance. Extensive experiments demonstrate that while a simple CNN achieves a maximum accuracy of 44% without augmentation, integration of VGG16 architecture substantially increases this figure to 84%. However, with our proposed augmentation techniques, accuracy further improves, reaching 60% for a simple CNN and an impressive 95% for a CNN combined with VGG16, marking a significant improvement over existing methods. Our approach sets a new benchmark for Caltech101 with a reported accuracy of 95% and shows promising adaptability to other imbalanced image datasets, offering a versatile solution to a pervasive challenge in image classification tasks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Towards Better Accuracy on Imbalance Image Datasets Using Image Augmentation and Convolutional Neural Networks

  • Sajid Ahmed,
  • Noriaki Yoshiura,
  • Saif Hassan,
  • Adil Khan

摘要

Imbalanced datasets present a significant challenge in supervised learning, particularly in image classification tasks where the number of instances varies widely among different classes, leading to reduced accuracy and bias toward more prevalent classes. This paper addresses the challenge posed by imbalanced datasets, exemplified by Caltech101, which includes a diverse range of instances per class, from as few as 31 to as many as 800. We introduce a robust technique combining artificial data augmentation with a pre-trained VGG16 architecture to significantly enhance classification accuracy. Our approach employs image augmentation techniques, including rotation, flipping, and noise injection, as a preprocessing step to enrich the dataset and mitigate class imbalance. Extensive experiments demonstrate that while a simple CNN achieves a maximum accuracy of 44% without augmentation, integration of VGG16 architecture substantially increases this figure to 84%. However, with our proposed augmentation techniques, accuracy further improves, reaching 60% for a simple CNN and an impressive 95% for a CNN combined with VGG16, marking a significant improvement over existing methods. Our approach sets a new benchmark for Caltech101 with a reported accuracy of 95% and shows promising adaptability to other imbalanced image datasets, offering a versatile solution to a pervasive challenge in image classification tasks.