<p>Extractive summaries of legal judgments are invaluable due to their ability to preserve the original text, facilitating direct reference when required. Existing research on extractive summarization of Indian legal judgments predominantly employs supervised methodologies, which rely on extensive and often difficult-to-obtaining expert annotations. State-of-the-art methods that use role labels or rhetorical role labels are largely limited to supervised learning paradigms (Bhattacharya et al. in Identification of rhetorical roles of sentences in indian legal judgments, 2019). In this study, we propose an unsupervised approach, Unsupervised Role-Labelled Knapsack Summarizer (<b>URL KnapSum</b>), which eliminates the need for expert annotations while effectively scaling to larger datasets. Our methodology begins by clustering sentences to group them into thematic categories. An optimized selection of sentences is performed using a knapsack algorithm using the similarity scores and length constraints, ensuring cluster-level diversity and global coherence. We evaluate the URL KnapSum using ROUGE (Lin in ROUGE: a package for automatic evaluation of summaries. In: Text summarization branches out, Barcelona, Spain. Association for Computational Linguistics, pp 74–81, 2004) scores for supervised evaluation and Kendall’s tau (Puka in Kendall’s Tau. Springer, Berlin, Heidelberg, pp 713–715, 2011) and Spearman’s correlation (Sedgwick in BMJ: Br Med J 349:g7327, 2014) for unsupervised evaluation. Additionally, we introduce a novel, reference-free evaluation metric, the <b>Top-K Analysis Metric</b>, which benchmarks the algorithm by measuring consistency between document and summary similarities. The results demonstrate the superiority of URL KnapSum over existing supervised methods, highlighting its effectiveness and scalability.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Extractive summarization of Indian legal judgements using unsupervised role labelling

  • Shruti Shreyasi,
  • Ayan Bandyopadhyay,
  • Partha Pratim Chakrabarti

摘要

Extractive summaries of legal judgments are invaluable due to their ability to preserve the original text, facilitating direct reference when required. Existing research on extractive summarization of Indian legal judgments predominantly employs supervised methodologies, which rely on extensive and often difficult-to-obtaining expert annotations. State-of-the-art methods that use role labels or rhetorical role labels are largely limited to supervised learning paradigms (Bhattacharya et al. in Identification of rhetorical roles of sentences in indian legal judgments, 2019). In this study, we propose an unsupervised approach, Unsupervised Role-Labelled Knapsack Summarizer (URL KnapSum), which eliminates the need for expert annotations while effectively scaling to larger datasets. Our methodology begins by clustering sentences to group them into thematic categories. An optimized selection of sentences is performed using a knapsack algorithm using the similarity scores and length constraints, ensuring cluster-level diversity and global coherence. We evaluate the URL KnapSum using ROUGE (Lin in ROUGE: a package for automatic evaluation of summaries. In: Text summarization branches out, Barcelona, Spain. Association for Computational Linguistics, pp 74–81, 2004) scores for supervised evaluation and Kendall’s tau (Puka in Kendall’s Tau. Springer, Berlin, Heidelberg, pp 713–715, 2011) and Spearman’s correlation (Sedgwick in BMJ: Br Med J 349:g7327, 2014) for unsupervised evaluation. Additionally, we introduce a novel, reference-free evaluation metric, the Top-K Analysis Metric, which benchmarks the algorithm by measuring consistency between document and summary similarities. The results demonstrate the superiority of URL KnapSum over existing supervised methods, highlighting its effectiveness and scalability.