<p>In the real world, most of the data obtained are hybrid data. Fully labeling such hybrid data is often time-consuming, labor intensive, and costly, while focusing solely on labeled samples may lead to the loss of critical information due to the limited number of labeled data. Research on feature selection using partially labeled data offers significant advantages, such as reducing dependency on labeled data and improving learning efficiency and performance. To address this issue, a semi-supervised feature selection method capable of handling partially labeled data is proposed based on the conditional discrimination index. First, by analyzing the characteristics of hybrid data, the distances between objects in the feature space are constructed, leading to the information granules with respect to feature subsets. Subsequently, in the label space, the missing labels are filled with the set composed of all existing labels, forming new partially labeled hybrid data. Considering the label distances between objects, a new tolerance relation is established to derive decision classes. Based on the information granules and decision classes, the significance of feature subsets is characterized using the discrimination index method. Then, a feature selection algorithm is designed depending on the significance. Experiments conducted on twelve real-world partially labeled hybrid datasets demonstrate that the proposed algorithm outperforms several existing feature selection algorithms in terms of classification accuracy and F1 score. Additionally, to validate the robustness of the algorithm, varying degrees of perturbations are introduced and tested on six of these datasets. The experimental results show that the algorithm still exhibits high stability, proving its applicability and reliability in complex and noisy environments. Statistical analyses further validate these findings. Finally, to further validate the practicality of the algorithm, it is deployed on an actual production line in a factory for fault detection, and significant results are achieved in the experimental tests.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A novel conditional discrimination index approach for feature selection in partially labeled hybrid data

  • Jiali He,
  • Mideth B. Abisado

摘要

In the real world, most of the data obtained are hybrid data. Fully labeling such hybrid data is often time-consuming, labor intensive, and costly, while focusing solely on labeled samples may lead to the loss of critical information due to the limited number of labeled data. Research on feature selection using partially labeled data offers significant advantages, such as reducing dependency on labeled data and improving learning efficiency and performance. To address this issue, a semi-supervised feature selection method capable of handling partially labeled data is proposed based on the conditional discrimination index. First, by analyzing the characteristics of hybrid data, the distances between objects in the feature space are constructed, leading to the information granules with respect to feature subsets. Subsequently, in the label space, the missing labels are filled with the set composed of all existing labels, forming new partially labeled hybrid data. Considering the label distances between objects, a new tolerance relation is established to derive decision classes. Based on the information granules and decision classes, the significance of feature subsets is characterized using the discrimination index method. Then, a feature selection algorithm is designed depending on the significance. Experiments conducted on twelve real-world partially labeled hybrid datasets demonstrate that the proposed algorithm outperforms several existing feature selection algorithms in terms of classification accuracy and F1 score. Additionally, to validate the robustness of the algorithm, varying degrees of perturbations are introduced and tested on six of these datasets. The experimental results show that the algorithm still exhibits high stability, proving its applicability and reliability in complex and noisy environments. Statistical analyses further validate these findings. Finally, to further validate the practicality of the algorithm, it is deployed on an actual production line in a factory for fault detection, and significant results are achieved in the experimental tests.