Comparative Evaluation of Hybrid ARIMA-XGBoost and Machine Learning Models for SO2 Forecasting in Gaziantep City
摘要
Air pollution represents a major public health challenge, particularly in rapidly developing urban environments. Accurate forecasting of sudden fluctuations in air pollutant concentrations is therefore essential for effective air quality management and environmental decision-making. In this study, SO2 concentrations (µg/m3) were forecast using four categories of forecasting approaches: statistical forecasting models, including Auto-Regressive Integrated Moving Average (ARIMA), Exponential Smoothing (ETS), and Prophet; standard machine-learning models, including linear regression (LR), k-nearest neighbours (KNN), multilayer perceptron neural networks (MLP), and support vector regression (SVR); ensemble-learning models, including Random Forest (RF) and Extreme Gradient Boosting (XGBoost); and hybrid forecasting models, including ARIMA-Boost and Prophet-Boost. A comprehensive comparative framework integrating statistical, machine-learning, ensemble-learning, and hybrid forecasting approaches was developed and systematically evaluated. Daily SO2 concentration data collected between 01/01/2020 and 31/07/2024 were obtained from the official air quality monitoring station operated by the Ministry of Environment, Urbanization and Climate Change in Gaziantep Province, Türkiye. Model performance was evaluated using mean absolute error (MAE), mean absolute percentage error (MAPE), and root mean squared error (RMSE). Among the evaluated models, the hybrid ARIMA-Boost framework achieved the best overall forecasting performance, producing the lowest MAE (2.61) and RMSE (3.71) values, whereas the ARIMA model yielded the lowest MAPE value (26.51). The findings demonstrate that hybrid statistical-machine-learning forecasting frameworks can substantially improve SO2 prediction accuracy compared with conventional statistical methods and standalone machine-learning approaches.