错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Graph-Based Machine Learning Approaches for Pangenomics

  • Indika Kahanda,
  • Joann Mudge,
  • Buwani Manuweera,
  • Thiruvarangan Ramaraj,
  • Alan Cleary,
  • Brendan Mumey

摘要

Deciphering the relationship between genotype and phenotype is a crucial yet challenging step in genetic research. Genome-wide association studies (GWAS) allow for phenotypic prediction by connecting underlying diversity in gene frequencies to complex phenotypic traits. GWAS analysis plays an important role in many different disciplines, such as the identification of genetic risk factors for diseases, and the identification of genotype that forms the primary basis for unique observable traits. However, GWAS has limitations because single nucleotide polymorphisms (SNPs), which are the types of variations identified relative to a single reference genome, are generally used for the analysis as they are easy to identify. These limitations lead to bias and make it challenging to use GWAS for studies involving multiple related genomes. On the other hand, due to the advancement of sequencing technology over the years, DNA repositories are seeing exponential growth recently. Consequently, pangenomes, which are collections of genomes from multiple related species, are becoming ubiquitous. Since these pangenomes capture the entire collection of sequences of multiple related species, they can be used to determine the population variation independent of a reference. Here we describe our novel graph-based approach for genomic analysis that allows generalizing GWAS to pangenomics, i.e. pangenome-wide association studies (PWAS). Our approach combines a fast graph construction algorithm for identifying Frequented Regions from compressed de Bruijn graphs, which are fed as input features for machine learning models. We demonstrate the utility of this (reference-free) Frequented Regions approach by developing machine learning regression models capable of predicting yeast phenotypes with a performance on par or better than (reference-based) SNPs.