TSformer: A Transformer-Based Model Focusing Specifically on the Fusion of Temporal-Spatial Features for Traffic Forecasting
摘要
With the rapid development of Intelligent Transportation Systems (ITS), how to accurately obtain traffic data predictions has become a key challenge. In recent years, many neural networks with complex structures have been designed to meet this challenge; however, these models often use the method of extracting temporal and spatial features independently and then fusing them, which ignores the intrinsic interconnectivity of spatio-temporal features. To address this problem, we propose TSformer, a Transformer-based model specifically designed to extract important features in both time and space. Our innovation lies in integrating Transformer’s cross-attention mechanism to pre-learn the model in the temporal dimension in phase before the spatial features are extracted, seamlessly fusing temporal and spatial dimensional features. We refer to this novel approach as Temporal-Spatial Cross Attention Fusion (TSCAF), which improves the model’s ability to capture the intrinsic connection between temporal and spatial features. Our experimental results on six traffic prediction datasets show the state-of-the-art performance of the model.