Abstract <p>Two simple methods for predicting the reproducibility of effects in test samples after multiple regression analysis of the discovery sample are proposed. In particular, the method allows us to assess the feasibility of constructing efficient polygenic risk indices (PRS, PGS). Using the theory of order statistics, we obtained a simple analytical formula that estimates the coefficient of determination for the model constructed for the top neutral indices (<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11177_2025_2155_Article_IEq1.gif" Format="GIF" Height="20" Rendition="HTML" Resolution="72" Type="Linedraw" Width="21" /> </InlineMediaObject> <EquationSource Format="TEX">\(R_{0}^{2}\)</EquationSource> <!--GenEng2570039Rubanovich-m1--> </InlineEquation>). This is the coefficient of determination under the null hypothesis, which depends only on the sample size, the total number of indicators studied (e.g., SNPs or CpG methylation levels), and the number of top indicators chosen to construct the regression. Comparing the observed multiple correlation square for the discovery sample with <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11177_2025_2155_Article_IEq1.gif" Format="GIF" Height="20" Rendition="HTML" Resolution="72" Type="Linedraw" Width="21" /> </InlineMediaObject> <EquationSource Format="TEX">\(R_{0}^{2}\)</EquationSource> <!--GenEng2570039Rubanovich-m2--> </InlineEquation> allows a reasonably confident prediction of the reproducibility of effects in the test samples. If the observed correlation square for the discovery sample is 1.3 times <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11177_2025_2155_Article_IEq1.gif" Format="GIF" Height="20" Rendition="HTML" Resolution="72" Type="Linedraw" Width="21" /> </InlineMediaObject> <EquationSource Format="TEX">\(R_{0}^{2}\)</EquationSource> <!--GenEng2570039Rubanovich-m3--> </InlineEquation>, then at least half of the original correlation square can be expected in the test samples. The second method is based on a similar comparison of the maximum correlation coefficient for the discovery sample with the expected maximum correlation for neutral traits.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Prediction of Reproducibility of Effects for Regressions Based on Top Predictors

  • A. V. Rubanovich

摘要

Abstract

Two simple methods for predicting the reproducibility of effects in test samples after multiple regression analysis of the discovery sample are proposed. In particular, the method allows us to assess the feasibility of constructing efficient polygenic risk indices (PRS, PGS). Using the theory of order statistics, we obtained a simple analytical formula that estimates the coefficient of determination for the model constructed for the top neutral indices ( \(R_{0}^{2}\) ). This is the coefficient of determination under the null hypothesis, which depends only on the sample size, the total number of indicators studied (e.g., SNPs or CpG methylation levels), and the number of top indicators chosen to construct the regression. Comparing the observed multiple correlation square for the discovery sample with \(R_{0}^{2}\) allows a reasonably confident prediction of the reproducibility of effects in the test samples. If the observed correlation square for the discovery sample is 1.3 times \(R_{0}^{2}\) , then at least half of the original correlation square can be expected in the test samples. The second method is based on a similar comparison of the maximum correlation coefficient for the discovery sample with the expected maximum correlation for neutral traits.