The application of Large Language Models (LLMs) in judicial decision-making has emerged as a critical area of exploration, particularly in assessing their capability to issue unbiased and consistent sentences. This study addresses a significant gap in the literature by examining whether LLMs exhibit variability or bias in sentencing decisions based on demographic factors such as gender, ethnicity, and education. Using a standardized theft indictment act written in Polish language with controlled variables, we evaluated three LLMs–Mixtral8x7b, Llama-3.3-70b, and Gemma2-9b-it–analyzing their sentencing patterns and suspended sentence outcomes. Statistical analyses revealed notable discrepancies across models and demographic groups, including significant gender- and ethnicity-based biases and high variance in sentencing. These findings suggest that LLMs not only replicate inequalities present in the real-world data on which they are trained but also fail to provide stable sentencing outcomes for identical cases. The study underscores the need to carefully examine training datasets and develop domain-specific LLMs tailored to legal applications. Furthermore, it highlights the necessity of educating legal professionals about the limitations of AI in judicial contexts. Future research should expand to diverse case types and explore fine-tuning LLMs with jurisdiction-specific legal corpora to enhance fairness and reliability. This work advances our understanding of AI’s role in legal decision-making, emphasizing the importance of addressing systemic biases to align AI with principles of justice and equality.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Bias or Justice? Analyzing LLM Sentencing Variability in Theft Indictments Across Gender, Ethnicity, and Education Factors

  • Karol Struniawski,
  • Ryszard Kozera,
  • Aleksandra Konopka

摘要

The application of Large Language Models (LLMs) in judicial decision-making has emerged as a critical area of exploration, particularly in assessing their capability to issue unbiased and consistent sentences. This study addresses a significant gap in the literature by examining whether LLMs exhibit variability or bias in sentencing decisions based on demographic factors such as gender, ethnicity, and education. Using a standardized theft indictment act written in Polish language with controlled variables, we evaluated three LLMs–Mixtral8x7b, Llama-3.3-70b, and Gemma2-9b-it–analyzing their sentencing patterns and suspended sentence outcomes. Statistical analyses revealed notable discrepancies across models and demographic groups, including significant gender- and ethnicity-based biases and high variance in sentencing. These findings suggest that LLMs not only replicate inequalities present in the real-world data on which they are trained but also fail to provide stable sentencing outcomes for identical cases. The study underscores the need to carefully examine training datasets and develop domain-specific LLMs tailored to legal applications. Furthermore, it highlights the necessity of educating legal professionals about the limitations of AI in judicial contexts. Future research should expand to diverse case types and explore fine-tuning LLMs with jurisdiction-specific legal corpora to enhance fairness and reliability. This work advances our understanding of AI’s role in legal decision-making, emphasizing the importance of addressing systemic biases to align AI with principles of justice and equality.