<p>While advanced Large Language Models (LLMs) can simulate human-like prosocial behaviors, the degree to which they align with human prosocial values and the underlying affective mechanisms remain unclear. This study addressed these gaps using the third-party punishment (TPP) paradigm, comparing LLM agents (GPT and DeepSeek series) with human participants (<i>n</i> = 100). The LLM agents (<i>n</i> = 500, 100 agents per model) were one-to-one constructed based on the demographic and psychological features of human participants. Prompt engineering was employed to initiate TPP games and record punitive decisions and affective responses in LLM agents. Results revealed that: (1) GPT-4o, DeepSeek-V3, and DeepSeek-R1 models demonstrated stronger fairness value alignment, choosing punitive options more frequently than humans in TPP games; (2) all LLMs replicated the human pathway from unfairness through negative affective response to punitive decisions, with stronger mediation effects of negative emotions observed in DeepSeek models than GPT models; (3) only DeepSeek-R1 exhibited the human-like positive feedback loop from previous punitive decisions to positive affective feedback and subsequent punitive choices; (4) most LLMs (excluding GPT-3.5) showed significant representational similarity to human affect-decision patterns; (5) notably, all LLMs displayed rigid affective dynamics, characterized by lower affective variability and higher affective inertia than the flexible, context-sensitive fluctuations observed in humans. These findings highlight notable advances in prosocial value alignment but underscore the necessity to enhance their affective dynamics to foster robust, adaptive prosocial LLMs. Such advancements could not only accelerate LLMs’ alignment with human values but also provide empirical support for the broader applicability of prosocial theories to LLM agents.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Prosocial behavior in Large Language Models: Value alignment and affective mechanisms

  • Hao Liu,
  • Yu Lei,
  • Zhen Wu

摘要

While advanced Large Language Models (LLMs) can simulate human-like prosocial behaviors, the degree to which they align with human prosocial values and the underlying affective mechanisms remain unclear. This study addressed these gaps using the third-party punishment (TPP) paradigm, comparing LLM agents (GPT and DeepSeek series) with human participants (n = 100). The LLM agents (n = 500, 100 agents per model) were one-to-one constructed based on the demographic and psychological features of human participants. Prompt engineering was employed to initiate TPP games and record punitive decisions and affective responses in LLM agents. Results revealed that: (1) GPT-4o, DeepSeek-V3, and DeepSeek-R1 models demonstrated stronger fairness value alignment, choosing punitive options more frequently than humans in TPP games; (2) all LLMs replicated the human pathway from unfairness through negative affective response to punitive decisions, with stronger mediation effects of negative emotions observed in DeepSeek models than GPT models; (3) only DeepSeek-R1 exhibited the human-like positive feedback loop from previous punitive decisions to positive affective feedback and subsequent punitive choices; (4) most LLMs (excluding GPT-3.5) showed significant representational similarity to human affect-decision patterns; (5) notably, all LLMs displayed rigid affective dynamics, characterized by lower affective variability and higher affective inertia than the flexible, context-sensitive fluctuations observed in humans. These findings highlight notable advances in prosocial value alignment but underscore the necessity to enhance their affective dynamics to foster robust, adaptive prosocial LLMs. Such advancements could not only accelerate LLMs’ alignment with human values but also provide empirical support for the broader applicability of prosocial theories to LLM agents.