<p>Legal documents, such as court judgments and legal briefs, are complex, lengthy, and challenging to analyze manually. The increasing digitization of legal resources has created a growing need for automated NLP-based systems to streamline document review, rhetorical role (RR) labeling, and summarization. However, manual annotation of RR-labels is time-consuming, expensive, and requires domain expertise, making large-scale labeled datasets difficult to obtain. To reduce manual annotation efforts, we propose a model-assisted approach to generate a silver-standard dataset using a gold-standard dataset. We fine-tuned multiple pre-trained models, including DistillBERT, LegalBERT, Legal RoBERTa, ERNIE 2.0, and SCI-BERT, on the human-annotated BUILDNyAI dataset. Legal RoBERTa achieved the highest accuracy of 71.7%, which we used to create LegSegSC, a dataset of 5472 Indian Supreme Court judgments with 1.1 million automatically labeled sentences. Our contributions include: (1) fine-tuning multiple models for RR classification in legal texts, (2) achieving state-of-the-art performance with Legal RoBERTa, (3) introducing a silver-standard large-scale dataset for future research, and (4) demonstrating the feasibility of reducing manual annotation effort while maintaining high-quality labels.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LegSegSC: A Silver-Standard Rhetorical Role Labeled Dataset of Indian Supreme Court Judgments

  • Saloni Sharma,
  • Piyush Pratap Singh,
  • Anshul Verma

摘要

Legal documents, such as court judgments and legal briefs, are complex, lengthy, and challenging to analyze manually. The increasing digitization of legal resources has created a growing need for automated NLP-based systems to streamline document review, rhetorical role (RR) labeling, and summarization. However, manual annotation of RR-labels is time-consuming, expensive, and requires domain expertise, making large-scale labeled datasets difficult to obtain. To reduce manual annotation efforts, we propose a model-assisted approach to generate a silver-standard dataset using a gold-standard dataset. We fine-tuned multiple pre-trained models, including DistillBERT, LegalBERT, Legal RoBERTa, ERNIE 2.0, and SCI-BERT, on the human-annotated BUILDNyAI dataset. Legal RoBERTa achieved the highest accuracy of 71.7%, which we used to create LegSegSC, a dataset of 5472 Indian Supreme Court judgments with 1.1 million automatically labeled sentences. Our contributions include: (1) fine-tuning multiple models for RR classification in legal texts, (2) achieving state-of-the-art performance with Legal RoBERTa, (3) introducing a silver-standard large-scale dataset for future research, and (4) demonstrating the feasibility of reducing manual annotation effort while maintaining high-quality labels.