How do noise labels in granulated datasets influence tree-based classification and rule generation performance: an experimental examination
摘要
While previous research has highlighted the merits of dealing with noisy labels in granulated datasets for classification tasks, the impact on decision rule generation remains unexplored. This study empirically investigates the effect of noise label removal on classification performance (CA) and rule generation performance, as measured by generation rate (GR) and simplicity rate (SR). The experiment granulated 36 datasets using an unsupervised equal width interval technique, subsequently detected and removed noise labels with a 60% filtering acceptance level. A decision tree classifier was trained, tested, and evaluated on both noisy and clean datasets to compare performance. Main results are obtained showing that: (1) Paired t-tests showed significant differences for CA (t = −2.72, p = 0.013) and GR (t = 2.65, p = 0.015), but not for SR (t = −1.25, p = 0.225). (2) Among the 23 noisy granulated datasets used, the coefficient of correlation (CoCr) between noise rate and CA improvement shows strong positive correlation (CoCr = 0.7204) whereas GR and SR improvements are negligible (CoCr = −0.1351 and −0.1147, respectively), (3) For the datasets with low noise rate (< 6.00%), the CoCr weakened (CoCr = 0.4004) for improvement of CA, but strengthened (CoCr = 0.5200 and 0.5467) for improvement of GR and SR, (4) For the datasets with high noise rate (> 10.00%), the CoCr of CA improvement remained substantial (CoCr = 0.6601), but GR and SR improvements are unlikely noteworthy (CoCr = 0.0506 and 0.1984). The research uncovers the impact of noise labels on rule generation performance.