Enhancing the forecasting of soil organic carbon concentration with vertical variations in highly heterogeneous karst areas in China using machine learning algorithms
摘要
Soil carbon plays a vital role in mitigating global warming. However, the optimal prediction model for soil organic carbon concentration (SOCc) in highly heterogeneous regions and its parameters remain unclear.
MethodsWe employed 8265 soil samples collected from a representative karst watershed in Southwest China to develop multiple models for predicting SOCc. The models included random forest (RF), support vector regression (SVR), artificial neural networks (ANNs), generalized boosted regression trees (GBRT), and multiple linear regression (MLR). We evaluated the predictive performance of these models across three soil depth intervals (0–20 cm, 20–50 cm, and 50–100 cm) and quantified the influence of driving factors on SOCc distribution.
ResultsCompared with MLR, RF (0.46 < R2 < 0.55) and GBRT (0.45 < R2 < 0.55) achieved the highest accuracy. In addition to predictive performance, our results demonstrated depth-dependent shifts in key environmental controls, particularly soil depth, bulk density, and gravel content, which reflect variations in physical and biogeochemical processes along the soil profile. For instance, surface SOCc was negatively correlated with bulk density, whereas subsoil SOCc showed the opposite trend, likely due to compaction effects and reduced microbial turnover. Gravel content, often neglected, positively contributed to SOC retention by supporting microbial habitats and buffering soil moisture.
ConclusionOverall, this study highlights the value of region-specific machine learning applications in uncovering spatial carbon dynamics and deepens our understanding of the mechanisms driving organic carbon distribution in karst environments shaped by complex biotic–abiotic interactions, offering practical guidance for land use planning and ecological restoration.