Structuring Macau’s Criminal Court Judgments with Large Language Models: Methodological Innovations for Data Accuracy and Sample Selection Bias
摘要
Court judgments, particularly in multilingual or low-standardization jurisdictions, remain an underutilized resource for empirical legal studies. This study develops an automated pipeline for creating structured data from Macau’s criminal court judgment files—documents presented in both Chinese and Portuguese—by leveraging a modern large language model (LLM) to parse and extract key variables. Validation against human-coded benchmarks shows high agreement with manual annotations, underscoring its reliability. We also address selection bias arising from nonrandom publication practices and outline strategies, such as instrumental-variables approaches and double machine learning, to enhance the robustness of statistical inference based on our dataset. The findings highlight the feasibility of transforming unstructured court records into analytically rich structured data, advancing computational criminology and empirical legal studies in Macau and more broadly.