Generative AI, particularly Large Language Models (LLMs), holds significant promise for enhancing judicial tasks, especially in automating the generation of legal cases during the sentencing phase. This paper introduces SARA (System for Analysis and summaRization of legal Actions), an innovative method for abstractive and multi-document summarization of legal proceedings. SARA employs GPT-4o, trained exclusively through in-context learning, utilizing Chain of Density (CoD) and CO-STAR prompt engineering techniques. These methods, adapted from judicial procedural knowledge, significantly improve the quality of generated summaries. Traditional evaluation metrics reveal effective training strategies but highlight their limitations in assessing summary quality. Therefore, we propose a qualitative evaluation methodology based on expert-generated questionnaires, focusing on essential content inclusion and proper report structuring. Inspired by this methodology, we trained another GPT model for large-scale summary evaluation. Evaluations of summaries of fifteen first-degree court cases from the Court of Justice of the State of Ceará show a significant advantage of in-context learning with CoD, emphasizing the role of domain knowledge and report style.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SARA - A Generative AI for Legal Process Summarization Based on Chain of Density Prompt Engineering

  • Francisco das Chagas Jucá Bomfim,
  • Joao Araujo Monteiro Neto,
  • Gilson Bezerra Filho,
  • Vasco Furtado,
  • Vládia Pinheiro

摘要

Generative AI, particularly Large Language Models (LLMs), holds significant promise for enhancing judicial tasks, especially in automating the generation of legal cases during the sentencing phase. This paper introduces SARA (System for Analysis and summaRization of legal Actions), an innovative method for abstractive and multi-document summarization of legal proceedings. SARA employs GPT-4o, trained exclusively through in-context learning, utilizing Chain of Density (CoD) and CO-STAR prompt engineering techniques. These methods, adapted from judicial procedural knowledge, significantly improve the quality of generated summaries. Traditional evaluation metrics reveal effective training strategies but highlight their limitations in assessing summary quality. Therefore, we propose a qualitative evaluation methodology based on expert-generated questionnaires, focusing on essential content inclusion and proper report structuring. Inspired by this methodology, we trained another GPT model for large-scale summary evaluation. Evaluations of summaries of fifteen first-degree court cases from the Court of Justice of the State of Ceará show a significant advantage of in-context learning with CoD, emphasizing the role of domain knowledge and report style.