<p>Feature selection aims to select effective features from the original features while maintaining unchanged classification ability and is an important data preprocessing technique. Fuzzy <InlineEquation ID="IEq4"> <EquationSource Format="TEX">\(\beta \)</EquationSource> <EquationSource Format="MATHML"><math> <mi>β</mi> </math></EquationSource> </InlineEquation>-covering effectively addresses the uncertainty, continuity, and missing values by constructing neighborhood approximation operators and fuzzy coverage particles suitable for real-valued data, playing an important role in feature selection for big data. In this paper, the fuzzy <InlineEquation ID="IEq5"> <EquationSource Format="TEX">\(\beta \)</EquationSource> <EquationSource Format="MATHML"><math> <mi>β</mi> </math></EquationSource> </InlineEquation>-covering theory is applied to model the given dataset, resulting in a fuzzy <InlineEquation ID="IEq6"> <EquationSource Format="TEX">\(\beta \)</EquationSource> <EquationSource Format="MATHML"><math> <mi>β</mi> </math></EquationSource> </InlineEquation>-covering incomplete real-valued decision information system (F<InlineEquation ID="IEq7"> <EquationSource Format="TEX">\(\beta \)</EquationSource> <EquationSource Format="MATHML"><math> <mi>β</mi> </math></EquationSource> </InlineEquation>-RVDIS). When applied to specific datasets, a distance function suitable for real-valued data is introduced to calculate the distance between the information values of objects in the information system. The fuzzy approximate relationship between objects is then obtained from the distance between information values. Finally, in the F<InlineEquation ID="IEq8"> <EquationSource Format="TEX">\(\beta \)</EquationSource> <EquationSource Format="MATHML"><math> <mi>β</mi> </math></EquationSource> </InlineEquation>-RVDIS, conditional information entropy, dependency degree, and importance measurement used to evaluate the feature subset are defined. Based on the characteristics of uncertainty measurement, an adaptive feature selection algorithm is designed, and numerical experiments and result analysis are conducted on the involved algorithms. Numerical experiments are performed on real datasets from the UCI machine learning repository, where the designed algorithm is compared with other advanced feature selection algorithms. Through comparisons in classification accuracy, feature selection rate, parameter analysis, and statistical tests, the effectiveness and superiority of the designed algorithm are demonstrated.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adaptive Feature Selection Based on Fuzzy \(\beta \)-Covering in Incomplete Real-Valued Decision Information System

  • Yemei Yang,
  • Pei Wang,
  • Qingguo Li

摘要

Feature selection aims to select effective features from the original features while maintaining unchanged classification ability and is an important data preprocessing technique. Fuzzy \(\beta \) β -covering effectively addresses the uncertainty, continuity, and missing values by constructing neighborhood approximation operators and fuzzy coverage particles suitable for real-valued data, playing an important role in feature selection for big data. In this paper, the fuzzy \(\beta \) β -covering theory is applied to model the given dataset, resulting in a fuzzy \(\beta \) β -covering incomplete real-valued decision information system (F \(\beta \) β -RVDIS). When applied to specific datasets, a distance function suitable for real-valued data is introduced to calculate the distance between the information values of objects in the information system. The fuzzy approximate relationship between objects is then obtained from the distance between information values. Finally, in the F \(\beta \) β -RVDIS, conditional information entropy, dependency degree, and importance measurement used to evaluate the feature subset are defined. Based on the characteristics of uncertainty measurement, an adaptive feature selection algorithm is designed, and numerical experiments and result analysis are conducted on the involved algorithms. Numerical experiments are performed on real datasets from the UCI machine learning repository, where the designed algorithm is compared with other advanced feature selection algorithms. Through comparisons in classification accuracy, feature selection rate, parameter analysis, and statistical tests, the effectiveness and superiority of the designed algorithm are demonstrated.