<p>In machine learning for QSAR/QSPR, the choice of train–test splitting algorithm alters both a model’s realized performance and the accuracy with which that performance is estimated from the held-out test set. Prior work has established that structure-aware splits such as Kennard–Stone and SPXY produce optimistically biased internal estimates, but these characterizations typically examine one method family at a time. Here we quantify how fifteen splitting strategies from four method families differ across fifteen drug-discovery datasets on two criteria (realized external benchmark performance and performance-estimation bias), and further assess whether splitting strategy affects model-selection quality when multiple model classes are optimized simultaneously. A single, consistent pattern holds across all fifteen datasets and all eleven evaluation measures (RMSE, MAE, MedAE, <InlineEquation ID="IEq1"><EquationSource Format="TEX">\(\hbox {R}^{2}\)</EquationSource><EquationSource Format="MATHML"><math><msup><mtext>R</mtext><mn>2</mn></msup></math></EquationSource></InlineEquation>, Spearman <InlineEquation ID="IEq2"><EquationSource Format="TEX">\(\rho\)</EquationSource><EquationSource Format="MATHML"><math><mi>ρ</mi></math></EquationSource></InlineEquation>, Pearson r, Kendall <InlineEquation ID="IEq3"><EquationSource Format="TEX">\(\tau\)</EquationSource><EquationSource Format="MATHML"><math><mi>τ</mi></math></EquationSource></InlineEquation>, enrichment factor at 5/10/20%, and BEDROC): Kennard–Stone-type methods (KS, SPXY, MDKS, Morais) achieve the strongest external benchmark performance but produce the most optimistically biased internal estimates (median Spearman bias <InlineEquation ID="IEq4"><EquationSource Format="TEX">\(\Delta \rho\)</EquationSource><EquationSource Format="MATHML"><math><mrow><mi mathvariant="normal">Δ</mi><mi>ρ</mi></mrow></math></EquationSource></InlineEquation> <InlineEquation ID="IEq5"><EquationSource Format="TEX">\(\approx\)</EquationSource><EquationSource Format="MATHML"><math><mo>≈</mo></math></EquationSource></InlineEquation> +0.17 to +0.22, SPXY most optimistic; median <InlineEquation ID="IEq6"><EquationSource Format="TEX">\(\hbox {R}^{2}\)</EquationSource><EquationSource Format="MATHML"><math><msup><mtext>R</mtext><mn>2</mn></msup></math></EquationSource></InlineEquation> bias <InlineEquation ID="IEq7"><EquationSource Format="TEX">\(\Delta \hbox {R}^{2}\)</EquationSource><EquationSource Format="MATHML"><math><mrow><mi mathvariant="normal">Δ</mi><msup><mtext>R</mtext><mn>2</mn></msup></mrow></math></EquationSource></InlineEquation> <InlineEquation ID="IEq8"><EquationSource Format="TEX">\(\approx\)</EquationSource><EquationSource Format="MATHML"><math><mo>≈</mo></math></EquationSource></InlineEquation> +0.26 to +0.41), diversity-based methods (OptiSim, Maximum/Minimum Dissimilarity) achieve the weakest benchmark performance, and cluster-shuffle variants, particularly k-means- and k-medoids-shuffle, provide the best-calibrated internal estimates (near-zero bias) while remaining competitive on benchmark performance. For every measure the differences in estimation bias between splitters greatly exceed the differences in benchmark performance (for Spearman <InlineEquation ID="IEq9"><EquationSource Format="TEX">\(\rho\)</EquationSource><EquationSource Format="MATHML"><math><mi>ρ</mi></math></EquationSource></InlineEquation>, a <InlineEquation ID="IEq10"><EquationSource Format="TEX">\(\approx\)</EquationSource><EquationSource Format="MATHML"><math><mo>≈</mo></math></EquationSource></InlineEquation>0.20 bias gap versus a <InlineEquation ID="IEq11"><EquationSource Format="TEX">\(\approx\)</EquationSource><EquationSource Format="MATHML"><math><mo>≈</mo></math></EquationSource></InlineEquation>0.02 benchmark gap), and the two criteria are inversely ranked (rank correlation <InlineEquation ID="IEq12"><EquationSource Format="TEX">\(\rho\)</EquationSource><EquationSource Format="MATHML"><math><mi>ρ</mi></math></EquationSource></InlineEquation> <InlineEquation ID="IEq13"><EquationSource Format="TEX">\(\approx\)</EquationSource><EquationSource Format="MATHML"><math><mo>≈</mo></math></EquationSource></InlineEquation> −0.70 to −0.83). For model selection among competing model classes, in contrast, no method achieves statistical separation and effect sizes are small, indicating that splitter choice has no demonstrated impact on which model class is ultimately selected. Splitter choice in drug-discovery QSAR therefore primarily affects the reliability of performance evaluation rather than realized model generalization. Cluster-shuffle splitting provides near-zero estimation bias while maintaining competitive external performance, making it the recommended default for offline model evaluation, where calibration is the primary concern. Because no splitter separates on model-selection quality, it is a sound default for model selection as well. Practitioners who wish to explicitly optimize model selection may use KS-type splitting to identify the best model configuration, then re-split with cluster-shuffle to obtain a calibrated performance estimate.</p><p><b>Scientific contribution</b></p><p> We provide the first head-to-head comparison of fifteen train–test splitting algorithms from four method families on a common drug-discovery regression benchmark, showing that their effect on the credibility of internal performance estimates is roughly an order of magnitude larger than their effect on realized model performance. We further show that this optimism is general across all eleven evaluation measures, give it a geometric explanation in terms of train–test distributional distance, and separate the splitting problem into two tasks, performance estimation and model selection, that have different optimal choices. On this basis we recommend cluster-shuffle splitting as a calibrated default; whereas prior work described the overoptimism of a single method family (Kennard–Stone), our comparison spans four families and all eleven measures and turns this into a concrete recommendation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Train–test splits matter more for evaluation than for performance

  • Davide Crucitti,
  • Adrián Mosquera Orgueira

摘要

In machine learning for QSAR/QSPR, the choice of train–test splitting algorithm alters both a model’s realized performance and the accuracy with which that performance is estimated from the held-out test set. Prior work has established that structure-aware splits such as Kennard–Stone and SPXY produce optimistically biased internal estimates, but these characterizations typically examine one method family at a time. Here we quantify how fifteen splitting strategies from four method families differ across fifteen drug-discovery datasets on two criteria (realized external benchmark performance and performance-estimation bias), and further assess whether splitting strategy affects model-selection quality when multiple model classes are optimized simultaneously. A single, consistent pattern holds across all fifteen datasets and all eleven evaluation measures (RMSE, MAE, MedAE, \(\hbox {R}^{2}\)R2, Spearman \(\rho\)ρ, Pearson r, Kendall \(\tau\)τ, enrichment factor at 5/10/20%, and BEDROC): Kennard–Stone-type methods (KS, SPXY, MDKS, Morais) achieve the strongest external benchmark performance but produce the most optimistically biased internal estimates (median Spearman bias \(\Delta \rho\)Δρ \(\approx\) +0.17 to +0.22, SPXY most optimistic; median \(\hbox {R}^{2}\)R2 bias \(\Delta \hbox {R}^{2}\)ΔR2 \(\approx\) +0.26 to +0.41), diversity-based methods (OptiSim, Maximum/Minimum Dissimilarity) achieve the weakest benchmark performance, and cluster-shuffle variants, particularly k-means- and k-medoids-shuffle, provide the best-calibrated internal estimates (near-zero bias) while remaining competitive on benchmark performance. For every measure the differences in estimation bias between splitters greatly exceed the differences in benchmark performance (for Spearman \(\rho\)ρ, a \(\approx\)0.20 bias gap versus a \(\approx\)0.02 benchmark gap), and the two criteria are inversely ranked (rank correlation \(\rho\)ρ \(\approx\) −0.70 to −0.83). For model selection among competing model classes, in contrast, no method achieves statistical separation and effect sizes are small, indicating that splitter choice has no demonstrated impact on which model class is ultimately selected. Splitter choice in drug-discovery QSAR therefore primarily affects the reliability of performance evaluation rather than realized model generalization. Cluster-shuffle splitting provides near-zero estimation bias while maintaining competitive external performance, making it the recommended default for offline model evaluation, where calibration is the primary concern. Because no splitter separates on model-selection quality, it is a sound default for model selection as well. Practitioners who wish to explicitly optimize model selection may use KS-type splitting to identify the best model configuration, then re-split with cluster-shuffle to obtain a calibrated performance estimate.

Scientific contribution

We provide the first head-to-head comparison of fifteen train–test splitting algorithms from four method families on a common drug-discovery regression benchmark, showing that their effect on the credibility of internal performance estimates is roughly an order of magnitude larger than their effect on realized model performance. We further show that this optimism is general across all eleven evaluation measures, give it a geometric explanation in terms of train–test distributional distance, and separate the splitting problem into two tasks, performance estimation and model selection, that have different optimal choices. On this basis we recommend cluster-shuffle splitting as a calibrated default; whereas prior work described the overoptimism of a single method family (Kennard–Stone), our comparison spans four families and all eleven measures and turns this into a concrete recommendation.