Comparative Analysis of Indian Legal Text Documents Using Large Language Models
摘要
This paper compares the capabilities of large language models—ChatGPT, Google Gemini, Bing Copilot, Claude, and Cohere—to summarize Indian legal text documents. Large language models have been found to develop with deep learning architectures like transformers, and they have made a difference in natural language processing by making more complex tasks of text generation and comprehension possible. In the present paper, a dataset consisting of judgments, resolutions, and legal opinions is used for the evaluation of models, based on quantitative metrics such as F1 score, accuracy, and recall, which are supplemented by qualitative evaluations from legal experts. ChatGPT performed the best, with an accuracy of 0.87, an F1 score of 0.91, and a recall of 0.83, having the average human evaluator score of 8.5. Next came a very close Google Gemini. On the other hand, Bing Copilot, Claude, and Cohere showed relatively poor performance. All these results clearly state the effectiveness of LLMs on legal text summarization, and ChatGPT happens to be the most reliable model.