Machine learning-based land-use regression models for predicting carbon dioxide concentrations in San Francisco Bay area
摘要
Carbon dioxide (CO2) is a key driver of anthropogenic climate change and cities have been identified as major sources of emissions. Urbanization and land use change are associated with rising urban CO2 emissions, highlighting the need to study spatiotemporal trends in intraurban CO2 to inform sustainable city planning. This study investigates the use of land use regression (LUR) to predict intraurban CO2 concentrations, using data from the BEACO2N monitoring network in the San Francisco Bay Area. Additionally, LUR is compared to machine learning (ML) algorithms capable of capturing non-linear relationships, representing a two-fold novel contribution. Model performance is evaluated using reserved data from training sensors as well as unseen sensor locations. For training sensors, extreme gradient boosting (XGBoost) and a convolutional neural network (CNN) achieved the highest predictive accuracy (R²=0.58), outperforming traditional LUR (R²=0.34). XGBoost and CNN also outperformed traditional LUR for unseen sensor locations, accounting for up to 42% of the variability in observed CO2 concentrations. These models offer insight into urban land use and carbon dynamics, supporting more informed approaches to urban planning and decarbonization.