Using Machine Learning Methods to Explore the Effects of Environmental Variables on Biodiversity
摘要
Environmental factors, notably air pollutants and meteorological conditions, significantly impact biodiversity distribution and ecosystem dynamics. Among these, \(PM_{10}\) is identified as a major stressor, affecting plant physiological responses, reducing species abundance, and modifying soil microbial activity and nutrient cycling. These effects are often compounded when \(PM_{10}\) interacts with climatic conditions. A comprehensive understanding of these combined effects is crucial for regions like Apulia, where biodiversity is shaped by both climate variability and localized human activities. The region remains underexplored in biodiversity modeling, particularly regarding the use of machine learning methods with the Multilevel Biodiversity Index (MBI). This study aims to address this gap by a comparison of the performance of several Machine Learning (ML) models, including Random Forests (RFs), Support Vector Machines (SVMs), Extreme Gradient Boosting (XGBoost), Decision Trees (DTs), and Linear Regression (LR), for predicting MBI based on \(PM_{10}\) concentrations and meteorological variables. Among the models tested, DTs and XGBoost performed best, reliably predicting MBI based on various evaluation metrics.