Indian Annual Report Assessment Using Large Language Models
摘要
The growth of large language models in the field of natural language processing (NLP) has sparked a heightened level of interest among academics in the extraction of practical information from existing documents. Additionally, the utilization of deep learning techniques has accelerated the development of efficient models for text mining purposes. Investment sectors focused on stock selection can significantly enhance their comprehension by accessing substantial and pertinent data within the complete collection of annual reports, specifically pertaining to the qualitative characteristics of the underlying firm. In this study, the author made available a public dataset for Indian annual reports and assessed the performance of two pre-trained language models, namely Bert and Roberta, for the specific task at hand. Additionally, the models were subjected to further fine-tuning using annual reports. The researcher examined the concepts of zero-shot learning and few-shot learning in the context of categorization, employing sentence transformers as a means of investigation. The results obtained from the fine-tuned language models and subsequent diverse classifiers exhibit promise, as they offer potential for financial analysts to integrate these discoveries into the realms of portfolio management and investment decision-making.