<p>Large Language Models (LLMs) are increasingly being applied to time-series forecasting, giving rise to a class of models referred to as time-series LLMs. While these models achieve competitive predictive accuracy, their reliability and structural consistency remain insufficiently understood. These models may produce forecasts that are numerically accurate on average yet statistically or temporally implausible, a phenomenon referred to as hallucination. Unlike conventional forecasting errors, hallucinations represent deviations from underlying temporal dynamics that exceed expected volatility patterns. This paper presents a systematic investigation of hallucination in time-series LLM forecasting. We introduce a quantitative evaluation framework that complements traditional regression metrics with two reliability-oriented measures: perplexity (PP), which reflects predictive uncertainty, and hallucination rate (HR), which measures statistically significant deviations from ground truth. Experiments on widely used benchmark datasets (Electricity and ETT variants) reveal a critical trade-off between forecasting accuracy and reliability. In several settings, improvements in average error metrics do not correspond to improved reliability; models can maintain low mean absolute error (MAE) while exhibiting high HR. To mitigate this issue, we evaluate two strategies:&#xa0;data-centric preprocessing, whose effectiveness depends on dataset characteristics, and structured tokenization, which consistently reduces hallucination across the evaluated datasets.&#xa0;Sensitivity analysis over quantile thresholds confirms the robustness of hallucination trends and model rankings. These results demonstrate that conventional regression metrics alone are insufficient for evaluating time-series LLMs and highlight the need for reliability-focused diagnostics when deploying LLM-based forecasting systems in high-stakes domains.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hallucination in Time-series Large Language Models: An empirical lnvestigation and analysis of mitigation strategies

  • Shamsu Abdullahi,
  • Kamaluddeen Usman Danyaro,
  • Haruna Chiroma,
  • Tieng Wei Koh,
  • Abubakar Zakari,
  • Yusuf Aliyu

摘要

Large Language Models (LLMs) are increasingly being applied to time-series forecasting, giving rise to a class of models referred to as time-series LLMs. While these models achieve competitive predictive accuracy, their reliability and structural consistency remain insufficiently understood. These models may produce forecasts that are numerically accurate on average yet statistically or temporally implausible, a phenomenon referred to as hallucination. Unlike conventional forecasting errors, hallucinations represent deviations from underlying temporal dynamics that exceed expected volatility patterns. This paper presents a systematic investigation of hallucination in time-series LLM forecasting. We introduce a quantitative evaluation framework that complements traditional regression metrics with two reliability-oriented measures: perplexity (PP), which reflects predictive uncertainty, and hallucination rate (HR), which measures statistically significant deviations from ground truth. Experiments on widely used benchmark datasets (Electricity and ETT variants) reveal a critical trade-off between forecasting accuracy and reliability. In several settings, improvements in average error metrics do not correspond to improved reliability; models can maintain low mean absolute error (MAE) while exhibiting high HR. To mitigate this issue, we evaluate two strategies: data-centric preprocessing, whose effectiveness depends on dataset characteristics, and structured tokenization, which consistently reduces hallucination across the evaluated datasets. Sensitivity analysis over quantile thresholds confirms the robustness of hallucination trends and model rankings. These results demonstrate that conventional regression metrics alone are insufficient for evaluating time-series LLMs and highlight the need for reliability-focused diagnostics when deploying LLM-based forecasting systems in high-stakes domains.