错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

On the Distribution-Free Discrimination: (A New Method of Discrimination)

  • Mahmoud Eltehiwy,
  • Abu Bakr Abdulmotaal,
  • Noura Taha Abu-Elmagd

摘要

This paper introduces the Q-method, a novel distribution-free discriminant procedure for two-group classification that leverages a proximity-weighted Q-statistic to measure discriminatory power. Unlike traditional rank-based approaches, which often require retaining the training sample for ranking new observations, the Q-method derives fixed linear coefficients directly from the original predictor values (robustly scaled by quartile deviations) after using ranks solely to compute relative variable importance. This eliminates training data retention, avoids reranking predictions, and provides an objective cutoff balancing misclassification rates and balance. The method imposes minimal assumptions (predictor independence and non-negative rank correlation after possible sign adjustment) and yields interpretable standardized coefficients, clear variable rankings, and probabilistic outputs via logistic transformation. Asymptotic normality of the Q-statistic enables per-variable significance testing, with finite-sample performance validated through extensive simulations showing reliable type I error control for moderate sample sizes. Empirical evaluations on benchmark datasets (Iris Setosa vs. Versicolor, Wisconsin Diagnostic Breast Cancer) and additional real-world applications (German Credit, Credit Card Default) demonstrate that the Q-method matches or outperforms classical parametric methods (LDA, QDA, logistic regression) and rank-transformed alternatives under normality, while showing clear superiority in non-normal, heteroscedastic, and skewed conditions. Monte Carlo simulations further confirm robustness across varying covariance structures, sample sizes, and moderate departures from normality. The Q-method offers a practical, assumption-light alternative for scenarios where distributional assumptions are uncertain or violated, particularly in small-to-moderate samples and resource-constrained settings, though current limitations include restriction to two groups and continuous predictors.