Interpretable genetic programming with SHAP-guided multi-objective optimization for scientific impact modeling
摘要
In academic evaluation, identifying high-impact researchers using bibliometric indicators requires models that are not only accurate but also interpretable. This study introduces a novel SHAP-guided Genetic Programming (SHAP-GP) framework that evolves symbolic ranking expressions guided by accuracy, complexity, and SHAP-based interpretability. Using curated datasets spanning four academic domains Mathematics, Civil Engineering, Computer Science, and Neuroscience our multi-objective optimization approach simultaneously maximizes classification performance, minimizes symbolic complexity, and improves explanation stability through SHAP metrics. We encode 64 bibliometric indicators as terminal nodes and evolve closed-form expressions that rely on as few as three to five features. A surrogate model is trained for each symbolic candidate to compute SHAP values, enabling quantification of feature Compactness and Stability. Comparative evaluations against Decision Trees, Explainable Boosting Machines (EBM), and Symbolic Regressors (SR) confirm the framework’s ability to produce domain-aligned, interpretable expressions with competitive F1-scores. In addition, we benchmark against three established GP baselines such as a standard accuracy-only GP, a parsimony-penalized GP, and a multi-objective GP without SHAP demonstrating that SHAP-GP consistently matches or outperforms these baselines while producing more compact and stable symbolic rules. Paired t-tests and Wilcoxon signed-rank tests validate the statistical significance of observed improvements. The proposed method offers a transparent and generalizable alternative to black-box classifiers for researcher profiling. By integrating interpretable AI into evolutionary design, the SHAP-GP framework advances decision transparency in research policy, funding allocation, and academic recognition.