Copy number variation (CNV) has become a focal point in biomedical research due to its critical role in detecting structural changes within DNA sequences, which usually involve the duplication and deletion of specific DNA segments. This study aims to advance the field by introducing a supervised Machine Learning approach with high specificity and sensitivity for classifying CNVs based on their pathogenicity. We employed five machine learning classification methods (Random Forest, Decision Tree, Bagging Classifier, XGBoost, and Adaboost), training the models on data from the dbVar and ClinVar databases and then validating them with data from the ClinGen and DECIPHER databases that had not previously been seen. Notably, Random Forest outperformed the other algorithms, with the highest accuracy of 94.3% on the internal dataset and a commendable accuracy of 83.3% on the validation data. The Random Forest model's ability to classify CNVs based on pathogenicity establishes it as a valuable tool for biomedical researchers and clinicians, offering advantages over existing tools. The developed model has the potential to enhance our understanding of CNVs and their implications in various diseases.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AI-Powered Insights into Genomic Variants for Sustainable Healthcare: Clinical Interpretation of Human Copy Number Variants

  • Maryem Ouhmouk,
  • Manal Nhili,
  • Hassan Ait Benhassou,
  • Mounia Abik

摘要

Copy number variation (CNV) has become a focal point in biomedical research due to its critical role in detecting structural changes within DNA sequences, which usually involve the duplication and deletion of specific DNA segments. This study aims to advance the field by introducing a supervised Machine Learning approach with high specificity and sensitivity for classifying CNVs based on their pathogenicity. We employed five machine learning classification methods (Random Forest, Decision Tree, Bagging Classifier, XGBoost, and Adaboost), training the models on data from the dbVar and ClinVar databases and then validating them with data from the ClinGen and DECIPHER databases that had not previously been seen. Notably, Random Forest outperformed the other algorithms, with the highest accuracy of 94.3% on the internal dataset and a commendable accuracy of 83.3% on the validation data. The Random Forest model's ability to classify CNVs based on pathogenicity establishes it as a valuable tool for biomedical researchers and clinicians, offering advantages over existing tools. The developed model has the potential to enhance our understanding of CNVs and their implications in various diseases.