This paper systematically investigates how large language models (LLMs) encode moral reasoning across six moral dimensions: care, fairness, loyalty, authority, sanctity, and liberty. We propose a novel interpretability pipeline that combines differential activation analysis, automated neuron description, and ablation experiments to identify specialized neurons aligned with each moral dimension. Our curated dataset of 240 validated moral and immoral statement pairs guides this exploration and reveals that certain neurons consistently exhibit increased activation in response to morally aligned statements. Notably, the care and sanctity dimensions show the largest sets of specialized neurons, whereas fairness and loyalty show fewer. We further demonstrate that ablating these neurons can causally modulate ethical decision-making, supporting the presence of discrete sub-circuits that influence moral outputs. Our findings not only advance the theoretical understanding of moral reasoning in LLMs, but also highlight avenues for targeted interventions and alignment. (All materials from this paper, including code, data, and experimental results, are available at https://github.com/coairesearch/mapping_moral_reasoning.git .)

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Mapping Moral Reasoning Circuits: A Mechanistic Analysis of Ethical Decision-Making in Large Language Models

  • Sigurd Schacht,
  • Carsten Lanquillon

摘要

This paper systematically investigates how large language models (LLMs) encode moral reasoning across six moral dimensions: care, fairness, loyalty, authority, sanctity, and liberty. We propose a novel interpretability pipeline that combines differential activation analysis, automated neuron description, and ablation experiments to identify specialized neurons aligned with each moral dimension. Our curated dataset of 240 validated moral and immoral statement pairs guides this exploration and reveals that certain neurons consistently exhibit increased activation in response to morally aligned statements. Notably, the care and sanctity dimensions show the largest sets of specialized neurons, whereas fairness and loyalty show fewer. We further demonstrate that ablating these neurons can causally modulate ethical decision-making, supporting the presence of discrete sub-circuits that influence moral outputs. Our findings not only advance the theoretical understanding of moral reasoning in LLMs, but also highlight avenues for targeted interventions and alignment. (All materials from this paper, including code, data, and experimental results, are available at https://github.com/coairesearch/mapping_moral_reasoning.git .)