<p>This paper evaluates the performance of GPT-4o in annotating decisions of the United Nations Committee on Economic, Social, and Cultural Rights and compares these results with manual annotations by trained law students and senior (legal) scholars. GPT-4o achieves human-level accuracy in basic annotations, but struggles with recall in citation extraction, particularly for complex legal references. Human annotators, while more reliable in citation extraction, introduce formatting inconsistencies and occasional errors due to sloppiness. In contrast, GPT-4o maintains high precision, but suffers from variability across repeated prompts, raising concerns about reproducibility. Beyond accuracy, this study highlights cost-effectiveness as a key advantage of GPT-4o. The model significantly reduces annotation time and expenses compared to human annotators, who require post-processing and expert supervision. While GPT-4o produces structured output with fewer formatting inconsistencies, its omissions and inconsistencies require human oversight. These findings highlight trade-offs in expertise, cost, and reliability between human and AI-driven annotation. Although GPT-4o is a viable tool for basic legal annotations, improvements in recall and consistency are needed for more complex tasks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The price of automated case law annotation: comparing the cost and performance of GPT-4o and student annotators

  • Iris Schepers,
  • Michelle Bruijn,
  • Martijn Wieling,
  • Michel Vols

摘要

This paper evaluates the performance of GPT-4o in annotating decisions of the United Nations Committee on Economic, Social, and Cultural Rights and compares these results with manual annotations by trained law students and senior (legal) scholars. GPT-4o achieves human-level accuracy in basic annotations, but struggles with recall in citation extraction, particularly for complex legal references. Human annotators, while more reliable in citation extraction, introduce formatting inconsistencies and occasional errors due to sloppiness. In contrast, GPT-4o maintains high precision, but suffers from variability across repeated prompts, raising concerns about reproducibility. Beyond accuracy, this study highlights cost-effectiveness as a key advantage of GPT-4o. The model significantly reduces annotation time and expenses compared to human annotators, who require post-processing and expert supervision. While GPT-4o produces structured output with fewer formatting inconsistencies, its omissions and inconsistencies require human oversight. These findings highlight trade-offs in expertise, cost, and reliability between human and AI-driven annotation. Although GPT-4o is a viable tool for basic legal annotations, improvements in recall and consistency are needed for more complex tasks.