Background <p>Time series models are extensively used for predicting infectious disease incidence through direct time series forecasting, and the Baidu Search Index (BSI) has recently been employed as a reference for analyzing disease trends. However, for predicting the hepatitis B incidence in China, it is essential to explore the most suitable models and forecasting methods, as well as the effectiveness of using BSI.</p> Methods <p>Data from the Chinese Center for Disease Control and Prevention, spanning January 2011 to December 2023, were used. Twelve models, including autoregressive integrated moving average (ARIMA), four deep learning models, and seven machine learning models were constructed to predict the reported incidence of hepatitis B in China. BSI-derived comprehensive search index (CSI) of hepatitis B-related keywords was employed as an exogenous variable for model construction. Both direct time-series forecasting and rolling forecasting were applied to predict the time series data in each model. Models’ predictive accuracy was evaluated using mean absolute percentage error (MAPE), weighted mean absolute percentage error (WMAPE), mean absolute error (MAE), mean square error (MSE), and root mean square error (RMSE), each accompanied by its 95% confidence interval (CI). Statistical significance was assessed using one-sided Diebold-Mariano tests.</p> Results <p>The ARIMA(0,1,1)(0,1,1)<sub>12</sub> model (WMAPE = 0.0821, 95% CI 0.0512–0.1312) outperformed other models in predicting hepatitis B reported incidence. In the primary split with single-step rolling forecasting, incorporating the CSI improved point-estimate predictive performance for nearly all models forecasting hepatitis B reported incidence, particularly in ARIMAX(1,1,1)(0,1,2)<sub>12</sub> (WMAPE 0.0548 vs. 0.0821), CNN (WMAPE 0.0845 vs. 0.1105), LSTM (WMAPE 0.0947 vs. 0.1067), Decision Tree (WMAPE 0.0972 vs. 0.1178), Random Forest (WMAPE 0.0956 vs. 0.1044), AdaBoost (WMAPE 0.0973 vs. 0.1049), and XGBoost (WMAPE 0.0996 vs.0.1039) models. Among these, statistically significant improvements were observed for ARIMA, CNN, and Decision Tree (<i>P</i> &lt; 0.05).</p> Conclusions <p>This study provides evidence that ARIMA model demonstrates excellent predictive performance for forecasting hepatitis B incidence trends in China. Additionally, BSI could serve as an effective supplement to traditional surveillance systems by providing real-time contemporaneous information that enhances situational awareness. Furthermore, adopting rolling forecasting might improve predictive accuracy for certain models and in specific contexts.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Predicting hepatitis B incidence trends in China using Baidu Search Index and time series models: a comparative study of rolling forecasting and direct forecasting strategies

  • Xing Yang,
  • Huanzhuo Mai,
  • Liuyan Lan,
  • Wudi Wei,
  • Yulan Xie,
  • Tingyan Luo,
  • Zhuoxin Li,
  • Wei Shi,
  • Jiegang Huang,
  • Langhuan Lei

摘要

Background

Time series models are extensively used for predicting infectious disease incidence through direct time series forecasting, and the Baidu Search Index (BSI) has recently been employed as a reference for analyzing disease trends. However, for predicting the hepatitis B incidence in China, it is essential to explore the most suitable models and forecasting methods, as well as the effectiveness of using BSI.

Methods

Data from the Chinese Center for Disease Control and Prevention, spanning January 2011 to December 2023, were used. Twelve models, including autoregressive integrated moving average (ARIMA), four deep learning models, and seven machine learning models were constructed to predict the reported incidence of hepatitis B in China. BSI-derived comprehensive search index (CSI) of hepatitis B-related keywords was employed as an exogenous variable for model construction. Both direct time-series forecasting and rolling forecasting were applied to predict the time series data in each model. Models’ predictive accuracy was evaluated using mean absolute percentage error (MAPE), weighted mean absolute percentage error (WMAPE), mean absolute error (MAE), mean square error (MSE), and root mean square error (RMSE), each accompanied by its 95% confidence interval (CI). Statistical significance was assessed using one-sided Diebold-Mariano tests.

Results

The ARIMA(0,1,1)(0,1,1)12 model (WMAPE = 0.0821, 95% CI 0.0512–0.1312) outperformed other models in predicting hepatitis B reported incidence. In the primary split with single-step rolling forecasting, incorporating the CSI improved point-estimate predictive performance for nearly all models forecasting hepatitis B reported incidence, particularly in ARIMAX(1,1,1)(0,1,2)12 (WMAPE 0.0548 vs. 0.0821), CNN (WMAPE 0.0845 vs. 0.1105), LSTM (WMAPE 0.0947 vs. 0.1067), Decision Tree (WMAPE 0.0972 vs. 0.1178), Random Forest (WMAPE 0.0956 vs. 0.1044), AdaBoost (WMAPE 0.0973 vs. 0.1049), and XGBoost (WMAPE 0.0996 vs.0.1039) models. Among these, statistically significant improvements were observed for ARIMA, CNN, and Decision Tree (P < 0.05).

Conclusions

This study provides evidence that ARIMA model demonstrates excellent predictive performance for forecasting hepatitis B incidence trends in China. Additionally, BSI could serve as an effective supplement to traditional surveillance systems by providing real-time contemporaneous information that enhances situational awareness. Furthermore, adopting rolling forecasting might improve predictive accuracy for certain models and in specific contexts.