Multi-label learning has garnered significant attention and application in the domains of machine learning, data mining, and pattern recognition. However, existing multi-label learning algorithms often fall short in adequately capturing the complex dynamic correlations among labels. To address this limitation, this paper proposes an online multi-label stream feature selection algorithm based on neigh-borhood approximation error rate and label correlation. Initially, the similarity between labels is computed, and correlated labels are partitioned into multiple subsets. The strongly correlated label sets thus obtained serve as the label space for the subsequent algorithm. Subsequently, instance boundaries are introduced to granulate all instances. Within the multi-label neighborhood approximation system, neighborhoods are determined based on the average intervals of samples, and a method for calculating feature importance is provided, offering a quantitative basis for feature selection. Finally, a dynamic stream feature selection framework is designed to iteratively select a subset of features that minimize the approximation error, thereby enhancing the efficiency and accuracy of feature selection. Experimental results demonstrate that the proposed algorithm consistently outperforms five other multi-label feature selection algorithms across all evaluation metrics.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Online Multi-label Streaming Feature Selection Based on Neighborhood Approximation Error Rate and Label Correlation

  • Siping Pan,
  • Juda Zhong,
  • Yu Mao,
  • Yaojin Lin

摘要

Multi-label learning has garnered significant attention and application in the domains of machine learning, data mining, and pattern recognition. However, existing multi-label learning algorithms often fall short in adequately capturing the complex dynamic correlations among labels. To address this limitation, this paper proposes an online multi-label stream feature selection algorithm based on neigh-borhood approximation error rate and label correlation. Initially, the similarity between labels is computed, and correlated labels are partitioned into multiple subsets. The strongly correlated label sets thus obtained serve as the label space for the subsequent algorithm. Subsequently, instance boundaries are introduced to granulate all instances. Within the multi-label neighborhood approximation system, neighborhoods are determined based on the average intervals of samples, and a method for calculating feature importance is provided, offering a quantitative basis for feature selection. Finally, a dynamic stream feature selection framework is designed to iteratively select a subset of features that minimize the approximation error, thereby enhancing the efficiency and accuracy of feature selection. Experimental results demonstrate that the proposed algorithm consistently outperforms five other multi-label feature selection algorithms across all evaluation metrics.