Developing Ensemble Models for Predicting the Antigenic Evolution
摘要
Evaluation of antigenic similarity between various strains of a single pathogenic virus is a crucial aspect of vaccine production. The conventional approaches employed to quantify the degree of similarity are based on the wet lab experiments, which are labor- and time-intensive. Therefore, in the last decade, various computer-aided and mathematical models have been developed to assist in acquiring earlier knowledge on the antigenic characteristics of currently circulating viruses. In this paper, optimized machine learning classifiers are applied to construct ensemble models for predicting antigenic variants of foot-and-mouth disease (FMD) virus. We provide a comprehensive assessment for embedding techniques of genetic sequences according to the type of sequences, DNA and protein, as well as by alternating the type of amino acid alphabet. Models that are trained using reduced amino acid alphabets supply the generation of new feature sets, which lead to reduce model complexity and increase ensemble diversity. The suggested stacking ensembles are constructed based on two diversity measures, double fault and Q-statistic. We conduct the evaluation of models on a relatively large test dataset. To the best of our knowledge, the proposed ensemble achieves a superior accuracy of 0.917. The results indicate that our model has potential as an exploratory tool for modeling the antigenicity of the FMD virus.