The methyl-CpG-binding protein 2 (MECP2) gene mutation is known to cause Rett syndrome and other severe neurodevelopmental disorders. Accurate characterization of the pathogenicity of these genetic variants is essential to understand the disease mechanisms and to allow early diagnosis. This understanding enables targeted therapeutic strategies mitigating the neurological consequences of these mutations. Despite advancements in genomic sequencing, determining the functional impact of novel mutations remains a challenge due to the high costs and time requirements of wet-lab validation. Existing computational models lack sufficient biological context and often fail to generalize well across diverse mutations. To bridge this gap, we propose a machine learning-based classification system that incorporates genomic sequence context, molecular consequence, and variant frequency for mutation analysis. This study introduces an approach by integrating transition or transversion ratios, nucleotide context, and feature importance ranking to enhance classification accuracy. The proposed approach leverages algorithms, specifically Random Forest and XGBoost, to address classification challenges inherent in genetic variant datasets. The XGBoost and Random Forest achieved 97.5

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Genomic Variant Classification for MECP2 Mutation Pathogenicity

  • Sravya Sri Mallampalli,
  • Chandra Mohan Dasari

摘要

The methyl-CpG-binding protein 2 (MECP2) gene mutation is known to cause Rett syndrome and other severe neurodevelopmental disorders. Accurate characterization of the pathogenicity of these genetic variants is essential to understand the disease mechanisms and to allow early diagnosis. This understanding enables targeted therapeutic strategies mitigating the neurological consequences of these mutations. Despite advancements in genomic sequencing, determining the functional impact of novel mutations remains a challenge due to the high costs and time requirements of wet-lab validation. Existing computational models lack sufficient biological context and often fail to generalize well across diverse mutations. To bridge this gap, we propose a machine learning-based classification system that incorporates genomic sequence context, molecular consequence, and variant frequency for mutation analysis. This study introduces an approach by integrating transition or transversion ratios, nucleotide context, and feature importance ranking to enhance classification accuracy. The proposed approach leverages algorithms, specifically Random Forest and XGBoost, to address classification challenges inherent in genetic variant datasets. The XGBoost and Random Forest achieved 97.5