错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Inferring Disease Risk from Genetics Using Cohort-Level Analysis

  • Mary Regina Boland

摘要

This chapter transitions readers, and budding health analysts, from analyzing a single participant’s direct-to-consumer (DTC) genetic results to investigating cohorts of DTC genetic data. We utilize our large cohort of 23andme OpenSNP users and link this larger cohort to information on phenotypic risk from their genetic results. We will utilize a simple example first based on predicting the blue eye color phenotype based on a few key single nucleotide polymorphisms (SNPs). This will allow us to explore the accuracy of key blue eye color SNPs from the literature. We will also evaluate potential confounders (e.g., color blindness of our participants) and how this could impact the self-reported phenotypes. Health analysts are often most concerned with disease risk prediction, and therefore, we will cover in detail an example of determining lactose intolerance based on genetic variants reported in the literature as being important in lactose intolerance. We will then evaluate those findings based on our individuals with self-reported lactose intolerance. We will explore regression models and generalized additive models (GAMs) and the need for spline terms to capture the non-linearity of certain confounder variables (height, weight, age). We will also cover imputation, which is very important when self-reported data is sparse, and the effect of imputation on model results. This chapter provides readers with an end-to-end analysis on the genetics of lactose intolerance that will be useful for health analysts as they explore other diseases/conditions/traits of interest. After reading this chapter, you should be able to confidently answer the following questions: