Current algorithms to estimate the pretest probability of coronary artery disease (CAD) are largely based on a set of well-defined clinical risk factors. However, scores based on these risk factors often have sub-optimal performance. Predictive models based on machine learning have been proposed to improve the accuracy of CAD risk prediction. These methods should ideally combine multimodal information from clinical, laboratory, and omics-based data streams. Graph representation learning methods are particularly promising, integrating data-driven approaches with prior biomedical knowledge. This paper investigates using PrimeKG, a knowledge graph proposed for precision medicine, to address the classification of Coronary Artery Stenosis (CAS) severity. We utilized clinical variables and whole blood transcriptome profiles, belonging to a cohort of adult patients without previous coronary events. We mapped these variables to the PrimeKG entities and adopted knowledge graph embedding to learn the graph representation. We represented each patient as a combination of the entity embeddings corresponding to the value of his/her feature in the dataset. We show that this knowledge-enriched patient representation is useful in the classification of CAD severity when used as input to predictive models, by improving the classification performance when compared to classic data-driven strategies in both single and multimodal settings.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Combining Clinical and Gene Expression Variables via Knowledge Graph Embedding for Prediction of Coronary Artery Stenosis

  • Giuseppe Albi,
  • Arianna Dagliati,
  • Chiara Vavassori,
  • Laura Pisani,
  • Mattia Chiesa,
  • Luca Piacentini,
  • Saima Mushtaq,
  • Gianluca Pontone,
  • Riccardo Bellazzi,
  • Gualtiero I. Colombo

摘要

Current algorithms to estimate the pretest probability of coronary artery disease (CAD) are largely based on a set of well-defined clinical risk factors. However, scores based on these risk factors often have sub-optimal performance. Predictive models based on machine learning have been proposed to improve the accuracy of CAD risk prediction. These methods should ideally combine multimodal information from clinical, laboratory, and omics-based data streams. Graph representation learning methods are particularly promising, integrating data-driven approaches with prior biomedical knowledge. This paper investigates using PrimeKG, a knowledge graph proposed for precision medicine, to address the classification of Coronary Artery Stenosis (CAS) severity. We utilized clinical variables and whole blood transcriptome profiles, belonging to a cohort of adult patients without previous coronary events. We mapped these variables to the PrimeKG entities and adopted knowledge graph embedding to learn the graph representation. We represented each patient as a combination of the entity embeddings corresponding to the value of his/her feature in the dataset. We show that this knowledge-enriched patient representation is useful in the classification of CAD severity when used as input to predictive models, by improving the classification performance when compared to classic data-driven strategies in both single and multimodal settings.