<b>Background</b> <p>Genomic selection relies on a variety of statistical and machine learning methods to predict phenotypes from genomic data. Since no single method consistently outperforms others across datasets, evaluating and comparing model performance is essential. However, standard evaluation metrics such as Pearson’s correlation coefficient and mean squared error treat genomic prediction as a regression problem, assessing overall fit rather than the effectiveness of selecting top-performing individuals for breeding. This disconnect can lead to suboptimal model selection in practice.</p> <b>Results</b> <p>To address this, we present the normalized cumulative gain (NCG) as an alternative evaluation measure that directly measures the phenotypic gain achieved from the individuals selected by the model. We applied this measure on four animal and plant datasets to compare nine commonly used methods for genomic prediction.</p> <b>Conclusions</b> <p>NCG offers an intuitive and interpretable measure of selection efficiency, focusing solely on the individuals that would actually be chosen. We further demonstrate that calculating the performance under all possible selection thresholds provides more information than a single or few arbitrary thresholds. This more granular analysis shows that the performance of the methods may differ under varying selection intensities and can provide guidance for the choice of selection intensity. Our approach is implemented in R and is available at <a href="https://github.com/FelixHeinrich/GS_Comparison_with_NCG">https://github.com/FelixHeinrich/GS_Comparison_with_NCG</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Normalized cumulative gain as an alternative evaluation measure for genomic selection models

  • Felix Heinrich,
  • Thomas M. Lange,
  • Faisal Ramzan,
  • Mehmet Gültas,
  • Armin O. Schmitt

摘要

Background

Genomic selection relies on a variety of statistical and machine learning methods to predict phenotypes from genomic data. Since no single method consistently outperforms others across datasets, evaluating and comparing model performance is essential. However, standard evaluation metrics such as Pearson’s correlation coefficient and mean squared error treat genomic prediction as a regression problem, assessing overall fit rather than the effectiveness of selecting top-performing individuals for breeding. This disconnect can lead to suboptimal model selection in practice.

Results

To address this, we present the normalized cumulative gain (NCG) as an alternative evaluation measure that directly measures the phenotypic gain achieved from the individuals selected by the model. We applied this measure on four animal and plant datasets to compare nine commonly used methods for genomic prediction.

Conclusions

NCG offers an intuitive and interpretable measure of selection efficiency, focusing solely on the individuals that would actually be chosen. We further demonstrate that calculating the performance under all possible selection thresholds provides more information than a single or few arbitrary thresholds. This more granular analysis shows that the performance of the methods may differ under varying selection intensities and can provide guidance for the choice of selection intensity. Our approach is implemented in R and is available at https://github.com/FelixHeinrich/GS_Comparison_with_NCG.