<p>Open-texture—e.g. vague, ambiguous, under-specified, or abstract terms—in regulatory documents lead to inconsistent interpretation, and are an obstacle to the automatic processing of regulation by computers. Identifying which parts of a legal text fall under open-texture is therefore a necessary requirement to make progress in automating the law. In this paper, we propose that large language models (LLMs) might provide an effective way to automatically detect open-texture in legal texts. We first investigate the obstacles by situating open-texture in the broader literature, and we test the hypothesis using two different LLMs—the proprietary gpt-3.5-turbo and the open-source llama-2-70b-chat—for the task of identifying open-texture in the General Data Protection Regulation. We evaluate their performance by asking 12 annotators to assess their output. We find, overall, that gpt-3.5-turbo overperforms llama-2-70b-chat on F<sub>1</sub>-scores (0.84 vs 0.67), and its high F<sub>1</sub>-score could make it a suitable alternative, or complement, to using human annotators. We also test the sensitivity of the findings against four further LLMs combined with six different prompts, and replicate a finding that there is low agreement between annotators when it comes to the identification of open-texture. We conclude the article by discussing the subjectivity of open-texture, the lessons to draw when testing for open-texture, and the consequences of using LLMs in the legal domain.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Identifying open-texture in regulations using LLMs

  • Clement Guitton,
  • Reto Gubelmann,
  • Ghassen Karray,
  • Simon Mayer,
  • Aurelia Tamò-Larrieux

摘要

Open-texture—e.g. vague, ambiguous, under-specified, or abstract terms—in regulatory documents lead to inconsistent interpretation, and are an obstacle to the automatic processing of regulation by computers. Identifying which parts of a legal text fall under open-texture is therefore a necessary requirement to make progress in automating the law. In this paper, we propose that large language models (LLMs) might provide an effective way to automatically detect open-texture in legal texts. We first investigate the obstacles by situating open-texture in the broader literature, and we test the hypothesis using two different LLMs—the proprietary gpt-3.5-turbo and the open-source llama-2-70b-chat—for the task of identifying open-texture in the General Data Protection Regulation. We evaluate their performance by asking 12 annotators to assess their output. We find, overall, that gpt-3.5-turbo overperforms llama-2-70b-chat on F1-scores (0.84 vs 0.67), and its high F1-score could make it a suitable alternative, or complement, to using human annotators. We also test the sensitivity of the findings against four further LLMs combined with six different prompts, and replicate a finding that there is low agreement between annotators when it comes to the identification of open-texture. We conclude the article by discussing the subjectivity of open-texture, the lessons to draw when testing for open-texture, and the consequences of using LLMs in the legal domain.