Learning from Multiple Noisy Annotations via Trustable Data Mixture
摘要
Our model-free approach utilizes Vicinal Risk Minimization (VRM) to address label noise in crowd-sourced datasets, avoiding complex adjustments based on Annotator-Specific Parameters (ASPs) that struggle with sparse data and identifiability issues. By selecting ‘trustable’ labels from the noisy dataset, we blend images across mislabeled classes in proportion to crowd-derived labels, enhancing input space with accurate class features and reducing discrepancies. This novel method applies VRM to crowd-based learning, operating effectively even with limited annotations and without assuming annotator independence. Its efficacy is confirmed through extensive testing across diverse datasets and confusion scenarios.