Noisy Instance Removal Using OWA-Based Fuzzy-Rough Sets
摘要
The reduction of the number of data instances is an important research area, particularly with a view to a reduction in the space requirements for lazy learning algorithms such as kNN. Previously, a fuzzy-rough prototype selection algorithm was proposed for this purpose, called OWAFRDC. This approach uses a criterion based on the upper and lower approximations of fuzzy-rough sets to assess the typicality of dataset instances. OWAFRDC was shown to preserve high quality instances and discard low quality instances. In this paper, a new instance quality criterion/measure is introduced to assess the quality of instances. The new criterion factors in the noisiness of instances in addition to their typicality. A numerical measure is calculated for each instance of a dataset based on the two mentioned criteria. The calculated values are used in the OWAFRDC algorithm to deliver condensed datasets. Non-parametric statistical tests show that the introduced quality measure improves the performance of OWAFRDC in terms of both accuracy and reduction rate.