Machine Learning-Based Tool for Efficient Discrimination Between Deleterious and Neutral Missense Mutations
摘要
We describe our new machine learning tool that we applied in the CAGI 6 experiment to predict whether single residue mutations in proteins are deleterious or benign. This was achieved by using single sequences only, without multiple sequence alignments or structural information. Instead, we used global characterizations of the protein sequence. Training and testing data of human mutations was obtained from ClinVar (ncbi.nlm.nih.gov/pub/ClinVar/). Testing was done on post-training data from ClinVar. This testing yielded high AUC and Matthews correlation coefficient (MCC). However, for genes with either sparse or unbalanced training data, the prediction accuracies are poor. The resulting prediction server is available online at http://www.mamiris.com/Shoni.cagi6 .