<p>Proteomic techniques now measure thousands of proteins circulating in blood at population scale, but successful translation into clinically useful protein biomarkers is hampered by our limited understanding of their origins. Here, we use machine learning to systematically identify a median of 20 factors (range: 1-37) out of &gt;1800 participant and sample charateristics that jointly explained an average of 19.4% (max. 100.0%) of the variance in plasma levels of ~3000 protein targets among 43,240 individuals. Proteins segregated into distinct clusters according to their explanatory factors, with modifiable characteristics explaining more variance compared to genetic variation (median: 10.0% vs 3.9%), and factors being largely consistent across the sexes and ancestral groups. We establish a knowledge graph that integrates our findings with genetic studies and drug characteristics to guide identification of potential drug target engagement markers. We demonstrate the value of our resource by identifying disease-specific biomarkers, like matrix metalloproteinase 12 for abdominal aortic aneurysm, and by developing a widely applicable framework for phenotype enrichment (R package: <a href="https://github.com/comp-med/r-prodente">https://github.com/comp-med/r-prodente</a>). All results are explorable via an interactive web portal (<a href="https://eur01.safelinks.protection.outlook.com/?url=https%3A%2F%2Fomicscience.org%2Fapps%2Fprotatlas&amp;data=05%7C02%7Cm.pietzner%40qmul.ac.uk%7C9d0796344b9e4dc5e1d408dd0fc05115%7C569df091b01340e386eebd9cb9e25814%7C0%7C0%7C638684040916065216%7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAuMDAwMCIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C&amp;sdata=QSEX6tbRJYfoDHqaOnSQeQgeUQrflaqd4slW4xrbvRE%3D&amp;reserved=0">https://omicscience.org/apps/prot_foundation</a>).</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Machine learning-guided deconvolution of plasma protein levels

  • Maik Pietzner,
  • Carl Beuchel,
  • Kamil Demircan,
  • Julian Hoffmann Anton,
  • Wenhuan Zeng,
  • Werner Römisch-Margl,
  • Summaira Yasmeen,
  • Burulça Uluvar,
  • Martijn Zoodsma,
  • Mine Koprulu,
  • Gabi Kastenmüller,
  • Julia Carrasco-Zanini,
  • Claudia Langenberg

摘要

Proteomic techniques now measure thousands of proteins circulating in blood at population scale, but successful translation into clinically useful protein biomarkers is hampered by our limited understanding of their origins. Here, we use machine learning to systematically identify a median of 20 factors (range: 1-37) out of >1800 participant and sample charateristics that jointly explained an average of 19.4% (max. 100.0%) of the variance in plasma levels of ~3000 protein targets among 43,240 individuals. Proteins segregated into distinct clusters according to their explanatory factors, with modifiable characteristics explaining more variance compared to genetic variation (median: 10.0% vs 3.9%), and factors being largely consistent across the sexes and ancestral groups. We establish a knowledge graph that integrates our findings with genetic studies and drug characteristics to guide identification of potential drug target engagement markers. We demonstrate the value of our resource by identifying disease-specific biomarkers, like matrix metalloproteinase 12 for abdominal aortic aneurysm, and by developing a widely applicable framework for phenotype enrichment (R package: https://github.com/comp-med/r-prodente). All results are explorable via an interactive web portal (https://omicscience.org/apps/prot_foundation).