Machine learning has achieved remarkable success in data classification tasks. However, due to their powerful feature learning capabilities, models may overfit to the label information in the training dataset, especially when trained on noisy labeled datasets that are common in real-world applications. This can lead to poor generalization performance. In this paper, we introduce a multi-stage ensemble approach with sample selection and label correction to build robust classification models under noisy labels. Our method employs optimized silhouette coefficient thresholds to distinguish between clean and noisy samples. By iteratively re-labeling noisy samples and updating the clean dataset at each stage, it enables models to learn diverse features and fully utilize the dataset. Compared to traditional sample selection methods, our approach achieves better generalization through the integration of models trained at each stage. Experiments on benchmark and real-world datasets demonstrate that our method significantly outperforms existing techniques in classifying datasets with noisy labels.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Novel Multi-stage Ensemble Method for Noisy Labels Using Sample Selection and Label Correction

  • Zhihong Yu,
  • Zongwen Fan,
  • Jin Gou

摘要

Machine learning has achieved remarkable success in data classification tasks. However, due to their powerful feature learning capabilities, models may overfit to the label information in the training dataset, especially when trained on noisy labeled datasets that are common in real-world applications. This can lead to poor generalization performance. In this paper, we introduce a multi-stage ensemble approach with sample selection and label correction to build robust classification models under noisy labels. Our method employs optimized silhouette coefficient thresholds to distinguish between clean and noisy samples. By iteratively re-labeling noisy samples and updating the clean dataset at each stage, it enables models to learn diverse features and fully utilize the dataset. Compared to traditional sample selection methods, our approach achieves better generalization through the integration of models trained at each stage. Experiments on benchmark and real-world datasets demonstrate that our method significantly outperforms existing techniques in classifying datasets with noisy labels.