This paper investigates the performance disparities between Small Language Models (SLMs) and Large Language Models (LLMs) in predicting stock price movements using data from two different datasets containing news articles and tweets. The study emphasizes the potential of SLMs as a more accessible and resource-efficient alternative to LLMs, enabling local and in-house deployment. Critical gaps are addressed, including the lack of direct price movement predictions, the utilization and comparison of State-of-the-Art (SotA) models, and the integration of diverse data sources. The research employed a fundamental trading strategy based on predicted stock price movement as the sole trading signal. The Phi-2 model, fine-tuned with Quantized Low-Rank Adaptation (QLoRA) on consumer-grade hardware, was compared with GPT-4, serving as a SotA benchmark. Performance was evaluated using accuracy, precision, recall, and F1-score. The results indicate that the fine-tuned SML (Phi-2) outperformed the LLM (GPT-4), albeit by a small margin, demonstrating the potential of a trained SML over a general LLM.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative Analysis and Evaluation of SLMs and LLMs for Stock Price Movement Prediction

  • Koen van der Leij,
  • George Manias,
  • Yik Kiu Leung,
  • Willem-Jan van den Heuvel

摘要

This paper investigates the performance disparities between Small Language Models (SLMs) and Large Language Models (LLMs) in predicting stock price movements using data from two different datasets containing news articles and tweets. The study emphasizes the potential of SLMs as a more accessible and resource-efficient alternative to LLMs, enabling local and in-house deployment. Critical gaps are addressed, including the lack of direct price movement predictions, the utilization and comparison of State-of-the-Art (SotA) models, and the integration of diverse data sources. The research employed a fundamental trading strategy based on predicted stock price movement as the sole trading signal. The Phi-2 model, fine-tuned with Quantized Low-Rank Adaptation (QLoRA) on consumer-grade hardware, was compared with GPT-4, serving as a SotA benchmark. Performance was evaluated using accuracy, precision, recall, and F1-score. The results indicate that the fine-tuned SML (Phi-2) outperformed the LLM (GPT-4), albeit by a small margin, demonstrating the potential of a trained SML over a general LLM.