<p>Accurate stock price prediction is crucial for informed financial decision-making and risk management in today’s volatile markets. In the evolution of the stock prediction methodologies, the trade-off between efficiency and accuracy has persisted as a critical challenge. Existing approaches based on conventional Transformers suffer from limited prediction accuracy and prolonged training times due to their self-attention mechanism and limited temporal information mining capabilities. To address these limitations, we propose a novel method called LAMFormer. This method integrates a multi-head Agent Attention mechanism to reduce the computational burden of traditional self-attention while incorporating a mixture-of-experts (MoE) module to adaptively extract multi-scale temporal features. Furthermore, the decoder replaces conventional self-attention layers with LSTM units, enhancing the model’s capacity to capture more diverse temporal characteristics&#xa0;and achieving efficient modeling of time-series features. We validate our approach on trading datasets from BYD, CATL, and CSSC, conducting extensive experiments across multiple prediction horizons. The results show that LAMFormer significantly outperforms baselines, achieving average improvements of 15.42% in MAE, 15.60% in RMSE, and 14.84% in MAPE, thereby demonstrating its superior performance and efficiency in stock price prediction.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Lamformer: LSTM-enhanced agent attention and mixture-of-experts transformer for efficient stock price prediction

  • Xiuyu Li,
  • Chenrui Bian,
  • Xiong Li,
  • Shiyuan Yu,
  • Botao Jiang

摘要

Accurate stock price prediction is crucial for informed financial decision-making and risk management in today’s volatile markets. In the evolution of the stock prediction methodologies, the trade-off between efficiency and accuracy has persisted as a critical challenge. Existing approaches based on conventional Transformers suffer from limited prediction accuracy and prolonged training times due to their self-attention mechanism and limited temporal information mining capabilities. To address these limitations, we propose a novel method called LAMFormer. This method integrates a multi-head Agent Attention mechanism to reduce the computational burden of traditional self-attention while incorporating a mixture-of-experts (MoE) module to adaptively extract multi-scale temporal features. Furthermore, the decoder replaces conventional self-attention layers with LSTM units, enhancing the model’s capacity to capture more diverse temporal characteristics and achieving efficient modeling of time-series features. We validate our approach on trading datasets from BYD, CATL, and CSSC, conducting extensive experiments across multiple prediction horizons. The results show that LAMFormer significantly outperforms baselines, achieving average improvements of 15.42% in MAE, 15.60% in RMSE, and 14.84% in MAPE, thereby demonstrating its superior performance and efficiency in stock price prediction.