Maxims for machines: operationalizing Kant’s universal law via inverse reinforcement learning
摘要
Recent work in AI ethics has revived Kantian deontology as a potential framework for AI alignment, yet questions remain about whether Kant’s formula of universal law (FUL) can provide concrete action guidance for artificial systems. This paper develops a Kantian-inspired procedural framework that defends the FUL as a viable alignment procedure by clarifying its normative structure and demonstrating how it can be operationalized computationally. Unlike standard reinforcement learning from human feedback, which learns preferences without imposing a prior formal constraint on permissibility, the paper proposes inverse reinforcement learning as a method for enabling AI systems to learn maxim formation and moral constraints from human exemplars, rather than relying on hard-coded rules. Through case studies in algorithmic hiring and content moderation, the paper shows that the FUL, understood as a procedural test, can offer principled and practically applicable guidance for AI alignment without presupposing machine moral agency.