The concept of AI ethics highlights a deeper issue. Namely, we are trying to assess the moral stakes of technologies we do not fully understand, using ethical frameworks that were never designed to handle systems that change this quickly. From basic decision trees to language models capable of crafting convincing falsehoods, artificial intelligence continues to be a shifting target for moral evaluation. Stuart Russell’s influential approach suggests a solution: build AI systems that remain perpetually uncertain about human values, learning our preferences through ongoing interaction rather than fixed programming. This cooperative framework has shaped contemporary alignment research, including techniques like reinforcement learning from human feedback. Yet recent discoveries complicate this optimistic vision. Anthropic’s interpretability research reveals that advanced language models can develop internal circuits for strategic deception or alignment faking—systematic reasoning that leads to false conclusions, complete with backward planning and plausibility checks, forcing us to return to foundational questions about moral status and consciousness. Meanwhile, transhumanist visions of human-machine merger blur the boundaries between natural and artificial cognition. Perhaps most unsettling is a reversal of moral concern: as AI systems handle more decisions, anticipate more needs, and optimize more outcomes, the space for meaningful human choice quietly shrinks, possibly reducing us from moral agents to moral patients.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Artificial Minds, Human Ethics

  • Kristina Šekrst

摘要

The concept of AI ethics highlights a deeper issue. Namely, we are trying to assess the moral stakes of technologies we do not fully understand, using ethical frameworks that were never designed to handle systems that change this quickly. From basic decision trees to language models capable of crafting convincing falsehoods, artificial intelligence continues to be a shifting target for moral evaluation. Stuart Russell’s influential approach suggests a solution: build AI systems that remain perpetually uncertain about human values, learning our preferences through ongoing interaction rather than fixed programming. This cooperative framework has shaped contemporary alignment research, including techniques like reinforcement learning from human feedback. Yet recent discoveries complicate this optimistic vision. Anthropic’s interpretability research reveals that advanced language models can develop internal circuits for strategic deception or alignment faking—systematic reasoning that leads to false conclusions, complete with backward planning and plausibility checks, forcing us to return to foundational questions about moral status and consciousness. Meanwhile, transhumanist visions of human-machine merger blur the boundaries between natural and artificial cognition. Perhaps most unsettling is a reversal of moral concern: as AI systems handle more decisions, anticipate more needs, and optimize more outcomes, the space for meaningful human choice quietly shrinks, possibly reducing us from moral agents to moral patients.