A Comparative Analysis of Machine-Learning Models for Discriminating Bacterial and Viral-Targeted Human Proteins
摘要
Bacterial and virally targeted human proteins are responsible for infectious diseases that are still prevalent. Understanding the difference between the bacterial protein and the viral protein smooths the way of diagnosing the disease. Machine learning models including RF, SVM, Decision Tree, k-NN, and two deep learning approaches, DNN and 1D CNN, can effectively provide fast distinction. For the human protein, the sequence, network, and gene ontology features have been considered, including GO IDs from UniProt. We have used amino acid composition (AAC), dipeptide composition (DC), pseudo-amino acid composition (PAAC), and composition-transition-distribution (CTD). The research focuses on the 1771 bacterial and 3916 viral human proteins to differentiate between bacterial and viral proteins. Our work shows 1D CNN scores higher in training (99.52%), testing (78.06%), and validation accuracy (88.86%), resulting in a preferable approach for deep learning for differentiating viral protein and bacterial protein.