<p>Feature selection is a pivotal step in machine learning, aimed at enhancing model performance by identifying the most relevant features while mitigating issues such as overfitting and computational complexity. This paper presents the Normalized Mean Difference (NMD), a novel, efficient univariate filter-based feature selection method designed to address limitations of traditional techniques. NMD quantifies a feature’s discriminative ability by computing the normalized difference between class means, ensuring scale invariance and robustness against feature magnitude biases. Key theoretical properties of NMD-symmetry, boundedness, and consistency-further establish its reliability across diverse datasets. The effectiveness of NMD is demonstrated through comprehensive experiments on benchmark datasets, including Heart Disease and Breast Cancer Wisconsin (Diagnostic), where it achieves competitive accuracy with fewer features compared to established methods such as Analysis of Variance (ANOVA), Chi-Squared (Ch-2), and Mutual Information (MI). Evaluations across multiple classifiers such as Decision Tree (DT), k-Nearest Neighbors (KNN), Logistic Regression (LR), Multilayer Perceptron (MLP), and Ridge Regression-underscore NMD’s versatility and broad applicability in tasks such as classification, regression, and clustering. With its theoretical grounding, simplicity, and practical scalability, NMD offers a robust feature selection framework. Implementations in Python and MATLAB are provided for reproducibility and practical deployment, accessible at <a href="https://drive.google.com/file/d/1113TU3H-s36GnKjOdHNbwgRwkbp011c3/view?usp=sharing">https://drive.google.com/file/d/1113TU3H-s36GnKjOdHNbwgRwkbp011c3/view?usp=sharing</a>this link.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Normalized mean difference (NMD): a novel filter-based feature selection method

  • Mohammed Mehdi Bouchene,
  • Mokeddem Fatima

摘要

Feature selection is a pivotal step in machine learning, aimed at enhancing model performance by identifying the most relevant features while mitigating issues such as overfitting and computational complexity. This paper presents the Normalized Mean Difference (NMD), a novel, efficient univariate filter-based feature selection method designed to address limitations of traditional techniques. NMD quantifies a feature’s discriminative ability by computing the normalized difference between class means, ensuring scale invariance and robustness against feature magnitude biases. Key theoretical properties of NMD-symmetry, boundedness, and consistency-further establish its reliability across diverse datasets. The effectiveness of NMD is demonstrated through comprehensive experiments on benchmark datasets, including Heart Disease and Breast Cancer Wisconsin (Diagnostic), where it achieves competitive accuracy with fewer features compared to established methods such as Analysis of Variance (ANOVA), Chi-Squared (Ch-2), and Mutual Information (MI). Evaluations across multiple classifiers such as Decision Tree (DT), k-Nearest Neighbors (KNN), Logistic Regression (LR), Multilayer Perceptron (MLP), and Ridge Regression-underscore NMD’s versatility and broad applicability in tasks such as classification, regression, and clustering. With its theoretical grounding, simplicity, and practical scalability, NMD offers a robust feature selection framework. Implementations in Python and MATLAB are provided for reproducibility and practical deployment, accessible at https://drive.google.com/file/d/1113TU3H-s36GnKjOdHNbwgRwkbp011c3/view?usp=sharingthis link.