Text Summarization and Sentiment Analysis Pipelines Using Large Language Models for Financial News
摘要
In an era dominated by vast amounts of unstructured web data, the need for sophisticated extraction and analysis methodologies is imperative. In this paper our goal is to address the challenge of summarizing financial news efficiently utilizing state-of-the-art transformer-based Large Language Models (LLMs) — specifically “t5-small”, and “sshleifer/distilbart-cnn-12–6” and conducting preliminary sentiment analysis. The performance of these summarization pipelines is rigorously evaluated against a comprehensive suite of metrics, including ROUGE, Keyword Overlap Score and Semantic Similarity Score for both models to validate the summaries’ accuracy and coherence. This investigation involved the CNBC news dataset, selecting 580 articles from the entire set available. By showcasing the capability of these models to distill pertinent information and hint at market sentiments using sentiment analysis pipelines, the research subtly underscores the emerging credibility and potential of AI techniques. This foundational work aims not only to enrich financial market analysis but also to catalyze further in-depth studies that refine Artificial Intelligence’s role in interpreting and leveraging the vast digital information landscape for strategic insights.