错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Language as a Lens: A Hybrid Text Summarization and Sentiment Analysis Approach for Multiclass Stock Return Prediction

  • Farshid Balaneji

摘要

This research explores the application of text summarization and sentiment analysis techniques in the multiclass classification of hourly stock price returns, studying six companies from the Dow Jones Index between 2017 and the first quarter of 2020. The study employs three distinct text summarization methods to efficiently process a substantial volume of financial news. This approach enhances the depth and accuracy of subsequent sentiment analysis. Sentiment is assessed through a combination of lexicon-based and deep learning methods. These insights, along with technical market data, are integrated into a Light Gradient Boosting Machine (LGBM) model, which is refined through Bayesian hyperparameter optimization and assessed via cross-validation. The results demonstrate that the model’s performance peaks at a standard deviation coefficient of 0.25, indicating an optimal balance for the three-class classification problem. Furthermore, the feature importance analysis reveals that while temporal and market-related factors play a significant role, sentiment features captured by FinBERT and VADER substantially contribute to the model’s predictive power. A notable finding of this research is the superior performance of FinBERT over GPT-3.5 in feature importance analysis, underscoring the efficacy of specialized language models in financial contexts. By marrying natural language processing techniques with machine learning, the study presents a novel approach to understanding the predictive power of news sentiment in financial markets.