Value-aligned but misguided: a dilemma in AI and AGI decision making
摘要
The development of artificial intelligence (AI) systems raises distinctive ethical and theoretical challenges not only because such systems will participate in human society in ways that invite moral appraisal, but also because a superintelligent agent is expected to exhibit a level of instrumental rationality that enables it to make decisions with social impact. This paper reframes the AI value alignment problem as a problem of robustness in decision making. Drawing on modified trolley-problem-style scenarios, influenced by the “Moral Machine” experiment, I argue that even AI systems governed by fixed ethical principles may produce actions that fail to align with those very principles under certain contextual conditions. The underlying issue lies in AI’s inability to reconcile the context-sensitive interpretation of values in a robust way. I suggest that this failure is best understood as arising from ambiguity in the interpretation of objectives. AI value alignment thus requires more than value specification; it demands the capacity to align values with contextually responsive belief formation and action selection.