Feature selection based on fuzzy rough fitting model with nominal distribution metric
摘要
Fuzzy rough sets constitute a significant granular computing model in knowledge discovery and have been widely applied to feature selection. However, in heterogeneous and nominal data, most existing fuzzy rough set methods rely on Hamming distance to measure dissimilarity between nominal attribute values, which fails to fully capture their underlying relationships and their impact on decision outcomes. To address this limitation, we propose a fuzzy rough fitting model with nominal distribution metric embedding (FR-NDM). First, the concept of nominal distribution and decision probability is defined, and suitable forms of the nominal distribution metric (NDM) are constructed for diverse data distribution scenarios. Second, a heterogeneous fuzzy information granule with dual-parameter adjustment is developed to accommodate complex data structures. Additionally, fitting approximation operators are established by introducing the judgment condition to ensure that samples attain the maximum membership degree within their respective decision categories. Third, a forward search feature selection algorithm is designed based on FR-NDM. Finally, the proposed method is evaluated on 24 public datasets and compared with 8 state-of-the-art feature selection methods. Experimental results demonstrate the superior performance of our approach.