Semi-supervised learning is a branch of machine learning that aims to build effective classifiers by leveraging a subset of the dataset accompanied by auxiliary information. This auxiliary information acts as supervision and guides the learning process. Common techniques for determining auxiliary information include Must-link/Cannot-link constraints, labels accompanying data points, predefined membership degrees, and more. Generally, the results often depend on the quality of the auxiliary information. Thus, varying the quality of the auxiliary information will yield different outcomes. The performance of the classifier may degrade if poorly chosen auxiliary information is used. The seed method requires a small set of data points in the sample space accompanied by labels. High-quality labels (good seeds) can enhance classification quality and minimize the number of queries from experts. Additionally, methods for selecting good labels (either automatically or from experts) during training are also subjects of research. These studies aim to minimize the number of queries while obtaining numerous high-quality labels. With this objective, we propose a new technique called FMN-Voting, tasked with gathering good seeds. FMN-Voting aims to (i) automatically determine auxiliary information for training neighboring data points, (ii) identify typical candidates for expert labeling, and (iii) assess the quality of the provided auxiliary information. To demonstrate the capability and effectiveness of FMN-Voting, experiments were conducted on several benchmark datasets. The results show the ability of FMN-Voting compared to several other methods with similar operations.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FMN-Voting: A Semi-supervised Clustering Technique Based on Voting Methods Using FMM Neural Networks

  • Dinh-Minh Vu,
  • Thanh-Son Nguyen,
  • Van Tinh Nguyen,
  • Duc-Luu Nguyen

摘要

Semi-supervised learning is a branch of machine learning that aims to build effective classifiers by leveraging a subset of the dataset accompanied by auxiliary information. This auxiliary information acts as supervision and guides the learning process. Common techniques for determining auxiliary information include Must-link/Cannot-link constraints, labels accompanying data points, predefined membership degrees, and more. Generally, the results often depend on the quality of the auxiliary information. Thus, varying the quality of the auxiliary information will yield different outcomes. The performance of the classifier may degrade if poorly chosen auxiliary information is used. The seed method requires a small set of data points in the sample space accompanied by labels. High-quality labels (good seeds) can enhance classification quality and minimize the number of queries from experts. Additionally, methods for selecting good labels (either automatically or from experts) during training are also subjects of research. These studies aim to minimize the number of queries while obtaining numerous high-quality labels. With this objective, we propose a new technique called FMN-Voting, tasked with gathering good seeds. FMN-Voting aims to (i) automatically determine auxiliary information for training neighboring data points, (ii) identify typical candidates for expert labeling, and (iii) assess the quality of the provided auxiliary information. To demonstrate the capability and effectiveness of FMN-Voting, experiments were conducted on several benchmark datasets. The results show the ability of FMN-Voting compared to several other methods with similar operations.