In the domain of time series forecasting, current Transformer-based models exhibit inadequate performance in aggregating dependencies both within and across variables, thereby resulting in suboptimal forecasting performance. Therefore, this paper proposes an improved attention mechanism and introduces an efficient forecasting model called Shareformer. The key points of this model are: (i) using the period of the time series data as the patch size to partition the sequence data; (ii) independently embedding each variable’s data; (iii) employing a shared attention mechanism that allocates the weights of intra-variable (inter-variable) attention scores to inter-variable (intra-variable) attention scores, thereby enabling information sharing between attentions. We conducted experiments on multiple datasets, and our model achieved the State-of-the-Art forecasting performance compared to other time series forecasting models. Additionally, we performed representation learning experiments, including transfer learning and self-supervised learning, and our model also outperformed self-supervised learning in these tasks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Shareformer: A Patch Transformer Model with Shared Attention for Multivariate Time Series Forecasting

  • Cunzhuo Lu,
  • Yeheng Jin,
  • Ziwei Zhang,
  • Xiangyu Peng,
  • Jing Chen

摘要

In the domain of time series forecasting, current Transformer-based models exhibit inadequate performance in aggregating dependencies both within and across variables, thereby resulting in suboptimal forecasting performance. Therefore, this paper proposes an improved attention mechanism and introduces an efficient forecasting model called Shareformer. The key points of this model are: (i) using the period of the time series data as the patch size to partition the sequence data; (ii) independently embedding each variable’s data; (iii) employing a shared attention mechanism that allocates the weights of intra-variable (inter-variable) attention scores to inter-variable (intra-variable) attention scores, thereby enabling information sharing between attentions. We conducted experiments on multiple datasets, and our model achieved the State-of-the-Art forecasting performance compared to other time series forecasting models. Additionally, we performed representation learning experiments, including transfer learning and self-supervised learning, and our model also outperformed self-supervised learning in these tasks.