Multi-objective reinforcement learning framework for beneficent artificial intelligence
摘要
As artificial intelligence (AI) systems are more widely deployed and utilized, they have greater potential to inflict harm or provide benefits to individuals. Designers must consider these concepts when building AI systems, but it is difficult to capture these notions in a mathematical framework. An ethical framework put forth by London and Heidari (Minds Mach, 2024. https://doi.org/10.1007/s11023-024-09696-8) provides structure and formal definitions of benefits and harms in relation to an individual’s life plans. This leaves an open question of how these concepts impact the decision-making of an AI system. We leverage their work to provide a direct translation of these theoretical ethical concepts to a standard multi-objective reinforcement learning (MORL) decision model. Using this model, we show that multi-objective rewards are necessary to capture benefits and harms in accordance with their definitions. We demonstrate how to capture these concepts in a MORL system and provide scenarios to highlight the utility of this work.