<p>Accurate multi-step air temperature forecasting in climatologically complex inland water environments remains a critical challenge for environmental monitoring and climate adaptation. This study benchmarks nine forecasting architectures (Naive Persistence, SARIMAX, Mamba, Kolmogorov–Arnold Networks, LightGBM, GRU, LSTM, TFT, and PatchTST) for multi-horizon hourly near-surface air temperature prediction over Lake Van, the world’s largest soda lake. Models were evaluated across five horizons (<i>t</i> + 1 to <i>t</i> + 168&#xa0;h) using a rolling-origin framework and five metrics (MAE, RMSE, sMAPE, NSE, <i>R</i><sup>2</sup>), with pairwise Diebold–Mariano (DM) tests to assess statistical significance. LightGBM achieved the lowest mean MAE and RMSE at all horizons (MAE: 1.078–3.001; NSE: 0.867–0.973), confirming the advantage of gradient-boosted ensembles on tabular environmental datasets. The results of the Diebold–Mariano (DM) tests indicate, however, that while LightGBM is superior in terms of average errors for each of the given standard metrics, it does not produce as consistently low an error variance as KAN did at all horizons. This result indicates that LightGBM’s superiority to the other models was based on both its average errors (point metric), which are often used to rank forecasts, and its lower variance, which is typically evaluated with a distributional test like the DM. Thus, these two different measures of forecast performance assess different, but equally important, aspects of how well a model can make predictions about future data. LightGBM minimizes average error while KAN delivers more consistently accurate predictions across the test period, demonstrating that relying on a single metric may lead to suboptimal model selection. Mamba demonstrated competitive short-horizon skill (NSE = 0.928 at <i>t</i> + 1) while remaining statistically comparable to LightGBM at extended lead times. LSTM substantially outperformed GRU, and PatchTST consistently outperformed TFT at all horizons. SARIMAX failed across all horizons (sMAPE &gt; 191%), confirming the inadequacy of linear models for non-linear sub-daily temperature dynamics. These findings advance the evidence base for data-driven environmental temperature forecasting and provide practical guidance for architecture selection across complex climatological applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-horizon hourly air temperature forecasting over Lake Van using classical, machine learning, and transformer-based models: a comparative benchmark with statistical significance testing

  • Mehmet Şamil Güneş

摘要

Accurate multi-step air temperature forecasting in climatologically complex inland water environments remains a critical challenge for environmental monitoring and climate adaptation. This study benchmarks nine forecasting architectures (Naive Persistence, SARIMAX, Mamba, Kolmogorov–Arnold Networks, LightGBM, GRU, LSTM, TFT, and PatchTST) for multi-horizon hourly near-surface air temperature prediction over Lake Van, the world’s largest soda lake. Models were evaluated across five horizons (t + 1 to t + 168 h) using a rolling-origin framework and five metrics (MAE, RMSE, sMAPE, NSE, R2), with pairwise Diebold–Mariano (DM) tests to assess statistical significance. LightGBM achieved the lowest mean MAE and RMSE at all horizons (MAE: 1.078–3.001; NSE: 0.867–0.973), confirming the advantage of gradient-boosted ensembles on tabular environmental datasets. The results of the Diebold–Mariano (DM) tests indicate, however, that while LightGBM is superior in terms of average errors for each of the given standard metrics, it does not produce as consistently low an error variance as KAN did at all horizons. This result indicates that LightGBM’s superiority to the other models was based on both its average errors (point metric), which are often used to rank forecasts, and its lower variance, which is typically evaluated with a distributional test like the DM. Thus, these two different measures of forecast performance assess different, but equally important, aspects of how well a model can make predictions about future data. LightGBM minimizes average error while KAN delivers more consistently accurate predictions across the test period, demonstrating that relying on a single metric may lead to suboptimal model selection. Mamba demonstrated competitive short-horizon skill (NSE = 0.928 at t + 1) while remaining statistically comparable to LightGBM at extended lead times. LSTM substantially outperformed GRU, and PatchTST consistently outperformed TFT at all horizons. SARIMAX failed across all horizons (sMAPE > 191%), confirming the inadequacy of linear models for non-linear sub-daily temperature dynamics. These findings advance the evidence base for data-driven environmental temperature forecasting and provide practical guidance for architecture selection across complex climatological applications.