Lag-length hyperparameter impacts on unit root testing: toward robust SARIMA integration order estimation
摘要
Accurate estimation of integration orders (d,D) is a foundational step in SARIMA modeling, directly influencing model stationarity, predictive performance, and structural stability. However, the sensitivity of these estimates to lag-length hyperparameterization within unit root testing procedures remains underexplored, posing a challenge for explainable and safe automation in time series forecasting. This study investigates the computational impact of various lag-selection strategies; specifically, AIC, BIC, HQIC, t-statistics, and rule-based heuristics, on the determination of differencing orders in the context of natural gas demand forecasting. Using high-resolution hourly consumption data from ten gas pressure reducing and metering stations (GPRMS) in Central Tunisia (2015–2023), we conduct extensive simulations across multiple temporal aggregations (weekly, monthly, and quarterly). Integration orders were derived using the augmented Dickey–Fuller (ADF) and Kwiatkowski–Phillips–Schmidt–Shin (KPSS) tests, applied under varying lag-length configurations. The resulting SARIMA models were benchmarked using MAPE and the overall weighted average (OWA) metrics. Results show lag-length selection critically impacts model performance, especially for low-frequency data. Monthly series favored BIC (52% of optimal configurations vs. 45% for AIC), with 8 of 10 stations achieving OWA < 1.0. Quarterly data showed clustered heterogeneity (BIC: 38%, AIC: 30%, Schwert: 24%) due to sample constraints. Weekly series exhibited lag-invariance, with 7 of 10 stations optimal across all methods. Default heuristics (Schwert’s rule) captured only 16–24% of optimal monthly/quarterly configurations, revealing a significant "automation gap" that undermines model reliability and safety. These findings highlight the critical role of lag-length specification in statistical preprocessing and provide practical recommendations for enhancing robustness in time series modeling workflows.