Kantian deontology for AI: alignment without moral agency
摘要
This paper explores the potential application of Kant’s moral philosophy to artificial intelligence (AI) and addresses two major objections. The first objection is that AI cannot fulfill Kant's standards for moral agency. I contend, however, that AI alignment with Kantian principles does not require moral agency in Kant's sense. I propose that the Categorical Imperative (CI) can serve as a useful framework for AI alignment, guiding the creation of maxims governing AI actions and testing their universalizability, particularly using the first principle of the CI which is the formula of the universal law (FUL). The second objection I address is the particularist critique to Kantian universalism, which is that Kantian universalism cannot tell us how to form maxims in a way that it allows sensitivity to context. I maintain that Kant’s framework can indeed accommodate context-sensitivity through practical judgment. But since AI are not the kinds of things to have practical judgment, I show that they have a functionally equivalent mechanism—transformer models—which can allow them form maxims that consider morally salient facts. Thus, supporting the claim that AI alignment is possible within a Kantian framework.