错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

BPSO-SLM: a binary particle swarm optimization-based self-labeled method for semi-supervised classification

  • Ruijuan Liu,
  • Junnan Li

摘要

The self-labeled methods have been favored by scholars in semi-supervised classification. Mislabeling is a great challenge for self-labeled methods and one of the reasons for mislabeling is that high-confidence unlabeled samples are found by mistake. While multiple variations of self-labeled methods have been developed, most existing strategies for finding high-confidence unlabeled samples heavily rely on specific assumptions. To solve the above issue, a binary particle swarm optimization-based self-labeled method (BPSO-SLM) is proposed and includes the following iterative self-labeled process: (a) A given classifier is trained on the set of labeled data; (b) The binary particle swarm optimization-based sample subspace optimization (BPSOSSO) is innovatively proposed to help BPSO-SLM find high-confidence unlabeled samples and low-confidence unlabeled samples from the set of unlabeled data; (c) The trained classifier is used to predict found high-confidence unlabeled samples; (d) High-confidence samples with pseudo labels are added to the set of labeled data, while low-confidence samples are returned to the set of unlabeled data and will be predicted again in the next iteration. The above process repeats until no high-confidence samples are found. After that, BPSO-SLM outputs the trained classifier during the iterative self-labeled process. The main characteristic of BPSO-SLM is that the strategy of finding high-confidence unlabeled samples makes little or no specific hypothesis about the geometric, distribution, and others of the selected high-confidence unlabeled samples. Intensive experiments on benchmark data prove that BPSO-SLM outperforms 8 state-of-the-art self-labeled methods in classification accuracy, Marco F-measure, and labeling error rate with various ratios of labeled data.