<p>The primary challenge in diagnosing cancer subtypes lies in the high dimensionality and sample imbalance of cancer gene expression data, which complicates accurate diagnosis from such complex datasets. To address the issues of feature selection (FS) and cancer subtype diagnosis for cancer gene expression data, a novel FS method with a double-layer parallel embedded structure has been proposed. This method is structured into two distinct stages: a preliminary screening stage followed by a fine screening stage. In the preliminary screening stage, a "mean-median" method was initially proposed to establish the threshold for the filter FS approach. During the fine screening stage, the Lasso regression algorithm was employed to refine the initial subset, thereby retaining what are referred to as "elite features." Both Sequential Forward Selection (SFS) and Sequential Backward Selection (SBS) algorithms preserved the ranking of these "elite features" along with their original total weight order. The enhanced SFS and SBS algorithms were then utilized to identify features that significantly influenced both the "elite features" and the initial subset. Subsequently, these two sets of features were merged and subjected to another round of screening. Based on the number of merged features, two distinct screening strategies were implemented. Ultimately, this process yielded a final set of selected features. Since the four filter methods were run separately in the preliminary screening stage, it can be regarded as a parallel operation of filter methods. The improved SFS and SBS algorithms were run separately in the fine screening stage, and can be regarded as a parallel operation of wrapper methods. Therefore, the proposed method is named PF-PSS. To demonstrate the effectiveness and superiority of the proposed PF-PSS FS method, it was evaluated on 20 cancer gene expression datasets. In the vast majority of cases, the proposed method demonstrates excellent performance in terms of classification accuracy and feature quantity selection. In particular, the PF-PSS FS method achieved a classification accuracy of up to 100% for 9 of the datasets, and the proportion of selected features for 12 of the datasets remained within 7%, with the minimum selected feature count even reaching 0.75%. On the other hand, the designed method has shown better performance in other evaluation criteria and running time.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

PF-PSS: a double-layer parallel embedded feature selection method for cancer gene expression data

  • Yu-Wei Song,
  • Jie-Sheng Wang,
  • Yu-Liang Qi,
  • Yu-Cai Wang,
  • Shi Li,
  • Hao-Ming Song,
  • Yi-Peng Shang-Guan

摘要

The primary challenge in diagnosing cancer subtypes lies in the high dimensionality and sample imbalance of cancer gene expression data, which complicates accurate diagnosis from such complex datasets. To address the issues of feature selection (FS) and cancer subtype diagnosis for cancer gene expression data, a novel FS method with a double-layer parallel embedded structure has been proposed. This method is structured into two distinct stages: a preliminary screening stage followed by a fine screening stage. In the preliminary screening stage, a "mean-median" method was initially proposed to establish the threshold for the filter FS approach. During the fine screening stage, the Lasso regression algorithm was employed to refine the initial subset, thereby retaining what are referred to as "elite features." Both Sequential Forward Selection (SFS) and Sequential Backward Selection (SBS) algorithms preserved the ranking of these "elite features" along with their original total weight order. The enhanced SFS and SBS algorithms were then utilized to identify features that significantly influenced both the "elite features" and the initial subset. Subsequently, these two sets of features were merged and subjected to another round of screening. Based on the number of merged features, two distinct screening strategies were implemented. Ultimately, this process yielded a final set of selected features. Since the four filter methods were run separately in the preliminary screening stage, it can be regarded as a parallel operation of filter methods. The improved SFS and SBS algorithms were run separately in the fine screening stage, and can be regarded as a parallel operation of wrapper methods. Therefore, the proposed method is named PF-PSS. To demonstrate the effectiveness and superiority of the proposed PF-PSS FS method, it was evaluated on 20 cancer gene expression datasets. In the vast majority of cases, the proposed method demonstrates excellent performance in terms of classification accuracy and feature quantity selection. In particular, the PF-PSS FS method achieved a classification accuracy of up to 100% for 9 of the datasets, and the proportion of selected features for 12 of the datasets remained within 7%, with the minimum selected feature count even reaching 0.75%. On the other hand, the designed method has shown better performance in other evaluation criteria and running time.