GPT vs human legal texts annotations: A comparative study with privacy policies
摘要
High-quality corpora of annotated privacy policies are scarce, yet essential for training, testing, and evaluating accurate machine learning and deep learning models. However, elaborating new privacy policy corpora remains an error-prone and resource-intensive task, heavily reliant on highly specialized and hard-to-find human annotators. Recent advancements in OpenAI’s Generative Pre-trained Transformers (GPTs) open the possibility of using them to annotate privacy policies with performance comparable to that of human annotators, thereby streamlining the process while reducing human resource demands. This paper presents a novel method for annotating privacy policies based on a codebook, a well-designed prompt, and analysis of logarithmic probabilities (logprobs) of GPT’s output tokens during the annotation process. We validated our method using the GPT-4o model and the well-known, open, multi-class, and multi-label OPP-115 corpus. GPT-4o achieved, without logprobs analysis, a performance comparable to eight out of the ten OPP-115 human annotators at the segment level of annotation, and to nine human annotators at the full-text level. Furthermore, incorporating logprobs analysis allowed GPT-4o to perform comparably to the ten OPP-115 human annotators at the full-text level of annotation, suggesting that context enhances the task. These findings demonstrate the potential of our proposed annotation method for creating privacy policy corpora with performance similar to that of human annotators while significantly reducing resource demands.
Graphical abstract