Development of a novel optimization modeling pipeline for range prediction of vectors with limited occurrence records in the Philippines: a bipartite approach
摘要
The upsurge in technical and epidemiological research employing Maximum Entropy (Maxent) establishes this machine-learning algorithm for species distribution modeling (SDM). Although Maxent robustly and accurately predicts the potential distribution of various species in different environments, data quality and varying hyperparameters influence its predictions. Optimizing hyperparameters can compensate for the rigidity of data quality. Addressing this caveat of Maxent, a bipartite approach (tuning and fine-tuning) in increasing model parsimony was developed to optimize the pipeline for range prediction of vectors with limited occurrence records in the Philippines. Tuned models reveal the influence of predictor collinearity on model accuracy, with a Pearson correlation threshold of 0.7 yielding the highest Area Under the Receiving Operator Characteristic Curve (AUC) score, analogous to popularly used methods in SDM. Fine-tuned models show that, contrary to the conventional pipeline, ΔAICc values approaching but not equal to zero produce a combination of hyperparameters (feature classes and regularization multiplier) leading to higher AUC scores. Fine-tuned models are more parsimonious and portray wider distributions than the a priori models generated using the default Maxent settings. This study integrates the best approaches to advance the conventional pipeline for Maxent modeling, substantiating the call for intensive surveying of vectors in a data-poor and high-burden country.