Identification of Mislabelled Images with Data Augmentation
摘要
Incorrectly labelled instances can significantly degrade both the performance of the machine learning model and the quality of its assessment. In this work, we describe a system for identifying mislabelled images. On the CIFAR-10 dataset, the achieved \(F_1\) score is higher by ten percentage points compared to another state-of-the-art result. Additionally, a new data augmentation scheme based on the selection of intensities for each transformation and the purity of an image is proposed and applied to the analyzed problem.