<p>We develop a “block” LASSO (blockLASSO) approach for training polygenic scores (PGS) and demonstrate its use in All of Us (AoU) and the UK Biobank (UKB). blockLASSO utilizes the approximate block diagonal structure (due to chromosomal partition of the genome) of linkage disequilibrium (LD). The new implementation can be used for exploratory and methods research where repeated PGS training is necessary and expensive. For 11 different phenotypes, in two different biobanks, and across 5 different ancestry groups (African, American, East Asian, European, and South Asian) – we demonstrate that blockLASSO is generally as effective for training PGS as a (global) LASSO. Previous work has shown penalized regression methods produce competitive PGS to alternative approaches. It has been shown that some phenotypes are more/less polygenic than others. Using sparse algorithms, an accurate PGS can be trained for type 1 diabetes (T1D) using <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="12864_2025_11505_Article_IEq1.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="39" /> </InlineMediaObject> <EquationSource Format="TEX">\({\sim }100\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mo>∼</mo> <mn>100</mn> </mrow> </math></EquationSource> </InlineEquation> single nucleotide variants (SNVs), but a PGS for body mass index (BMI) would need more than 10k SNVs. blockLASSO produces similar PGS for phenotypes while training with just a fraction of the variants per block. Within AoU (using only genetic information) block PGS for T1D reaches an AUC of <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="12864_2025_11505_Article_IEq2.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="60" /> </InlineMediaObject> <EquationSource Format="TEX">\(0.63_{\pm 0.02}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>0</mn> <mo>.</mo> <msub> <mn>63</mn> <mrow> <mo>±</mo> <mn>0.02</mn> </mrow> </msub> </mrow> </math></EquationSource> </InlineEquation> and for BMI a correlation of <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="12864_2025_11505_Article_IEq3.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="60" /> </InlineMediaObject> <EquationSource Format="TEX">\(0.21_{\pm 0.01}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>0</mn> <mo>.</mo> <msub> <mn>21</mn> <mrow> <mo>±</mo> <mn>0.01</mn> </mrow> </msub> </mrow> </math></EquationSource> </InlineEquation>, whereas a global LASSO approach which finds for T1D an AUC <InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="12864_2025_11505_Article_IEq4.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="60" /> </InlineMediaObject> <EquationSource Format="TEX">\(0.65_{\pm 0.03}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>0</mn> <mo>.</mo> <msub> <mn>65</mn> <mrow> <mo>±</mo> <mn>0.03</mn> </mrow> </msub> </mrow> </math></EquationSource> </InlineEquation> and BMI a correlation <InlineEquation ID="IEq5"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="12864_2025_11505_Article_IEq5.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="60" /> </InlineMediaObject> <EquationSource Format="TEX">\(0.19_{\pm 0.03}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>0</mn> <mo>.</mo> <msub> <mn>19</mn> <mrow> <mo>±</mo> <mn>0.03</mn> </mrow> </msub> </mrow> </math></EquationSource> </InlineEquation>. This new block approach is more computationally efficient and scalable than naive global machine learning approaches and makes it ideal for exploratory methods investigations based on penalized regression.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Efficient blockLASSO for polygenic scores with applications to All of Us and UK Biobank

  • Timothy G. Raben,
  • Louis Lello,
  • Erik Widen,
  • Stephen D. H. Hsu

摘要

We develop a “block” LASSO (blockLASSO) approach for training polygenic scores (PGS) and demonstrate its use in All of Us (AoU) and the UK Biobank (UKB). blockLASSO utilizes the approximate block diagonal structure (due to chromosomal partition of the genome) of linkage disequilibrium (LD). The new implementation can be used for exploratory and methods research where repeated PGS training is necessary and expensive. For 11 different phenotypes, in two different biobanks, and across 5 different ancestry groups (African, American, East Asian, European, and South Asian) – we demonstrate that blockLASSO is generally as effective for training PGS as a (global) LASSO. Previous work has shown penalized regression methods produce competitive PGS to alternative approaches. It has been shown that some phenotypes are more/less polygenic than others. Using sparse algorithms, an accurate PGS can be trained for type 1 diabetes (T1D) using \({\sim }100\) 100 single nucleotide variants (SNVs), but a PGS for body mass index (BMI) would need more than 10k SNVs. blockLASSO produces similar PGS for phenotypes while training with just a fraction of the variants per block. Within AoU (using only genetic information) block PGS for T1D reaches an AUC of \(0.63_{\pm 0.02}\) 0 . 63 ± 0.02 and for BMI a correlation of \(0.21_{\pm 0.01}\) 0 . 21 ± 0.01 , whereas a global LASSO approach which finds for T1D an AUC \(0.65_{\pm 0.03}\) 0 . 65 ± 0.03 and BMI a correlation \(0.19_{\pm 0.03}\) 0 . 19 ± 0.03 . This new block approach is more computationally efficient and scalable than naive global machine learning approaches and makes it ideal for exploratory methods investigations based on penalized regression.