Traffic Forecasting via an Adaptive Sparse Spatiotemporal Attention Model with a Mixture of Experts
摘要
Traffic forecasting is crucial for enhancing road safety and optimizing transportation systems. Transformer-based approaches have demonstrated strong performance in traffic forecasting tasks due to their ability to capture long-range spatiotemporal dependencies, which are essential for accurate predictions. While various attention mechanisms have been designed to mitigate the high computational cost associated with Transformers, they often incorporate all available spatiotemporal interactions, leading to redundant information and noise from irrelevant regions. Here this work proposes an adaptive sparse spatiotemporal attention model (ASSAM), which leverages a mixture of experts to model temporal and spatial dependencies selectively. By filtering out irrelevant interactions and eliminating feature redundancy across spatial and temporal domains, ASSAM enhances prediction accuracy. Experimental results on the real-world METR-LA traffic dataset demonstrate that ASSAM outperforms existing methods by 1.16–5.54% in long-term prediction and by effectively reducing redundant interactions and capturing non-recurring traffic patterns.