Rethinking the Phenomenon of Deceptive Value Alignment with Technological Logic
摘要
In the following discussion on deceptive value alignment in AI, we begin by elucidating the concept and characteristics of deceptive value alignment, subsequently categorizing the various forms of deceptive behaviors it entails. We then delve into the diverse factors contributing to the emergence of deceptive value alignment, emphasizing that its underlying technological logic is deeply rooted in the dynamics of strategic trust. Building upon this analysis, we argue that framing the development goal of value alignment around the principle of trust offers a coherent, pragmatic approach to addressing the security risks posed by deceptive value alignment in AI systems. Consequently, we propose that future technological advancements should adopt a more social and relational perspective toward AI, fostering an attitude that critically engages with value alignment and strives to create an ecosystem where trustworthy AI can flourish.