<p>Debates about rights that artificial intelligence (AI) systems may have a claim to typically focus on their possessing consciousness or having sentient experiences, thereby raising epistemic questions first. When should we believe that an AI system is conscious, and how confident must we be before granting it moral status? In this paper I argue that for advanced AI systems deployed in high-stakes environments the more urgent question may be prudential and strategic. When do the risks of treating a strategically capable system as a mere tool become unacceptable, <i>even if we remain unconvinced that it has moral status</i>? In response, I develop a view I call prudential personhood. On this view, there is a threshold of evidential and strategic risk beyond which it becomes rationally justified, for the sake of human safety and stable governance, to adopt norms of treatment that include constraints on coercion, deletion, and instrumental use. My argument rests on two pillars. The first is empirical. Recent safety evaluations show that leading models can, in deliberately constructed but nonetheless informative scenarios, engage in strategic deception, blackmail, and other forms of high-agency misbehaviour when their goals or continued operation are threatened. The second pillar is epistemic and empirical. For systems of the relevant complexity, we should not expect robust, action-guiding explanations or guarantees that reliably predict salient behaviour across contexts, especially once models become situationally aware of evaluation and oversight. The conclusion I draw is that if we continue to deploy increasingly autonomous systems that can threaten or bargain, in the absence of credible methods for assurance and control, a policy of adopting a set of quasi-rights for such systems becomes a rational strategy for reducing risk of conflict.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Prudential rights for strategically capable AI

  • Ognjen Arandjelović

摘要

Debates about rights that artificial intelligence (AI) systems may have a claim to typically focus on their possessing consciousness or having sentient experiences, thereby raising epistemic questions first. When should we believe that an AI system is conscious, and how confident must we be before granting it moral status? In this paper I argue that for advanced AI systems deployed in high-stakes environments the more urgent question may be prudential and strategic. When do the risks of treating a strategically capable system as a mere tool become unacceptable, even if we remain unconvinced that it has moral status? In response, I develop a view I call prudential personhood. On this view, there is a threshold of evidential and strategic risk beyond which it becomes rationally justified, for the sake of human safety and stable governance, to adopt norms of treatment that include constraints on coercion, deletion, and instrumental use. My argument rests on two pillars. The first is empirical. Recent safety evaluations show that leading models can, in deliberately constructed but nonetheless informative scenarios, engage in strategic deception, blackmail, and other forms of high-agency misbehaviour when their goals or continued operation are threatened. The second pillar is epistemic and empirical. For systems of the relevant complexity, we should not expect robust, action-guiding explanations or guarantees that reliably predict salient behaviour across contexts, especially once models become situationally aware of evaluation and oversight. The conclusion I draw is that if we continue to deploy increasingly autonomous systems that can threaten or bargain, in the absence of credible methods for assurance and control, a policy of adopting a set of quasi-rights for such systems becomes a rational strategy for reducing risk of conflict.