Reflections on the AI alignment problem
摘要
The Alignment Problem in artificial intelligence concerns how to insure that artificial general intelligence (AGI) conforms to human goals and values and remains under human control. The concept of general intelligence, modelled on human and animal behavior, lacks coherence. The ideal of autonomy inherent in AGI conflicts with the ideal of external control. Truly autonomous agents are necessarily embodied, but embodiment implies more than physical instantiation or sensory input. It means being an autopoietic system (like a natural organism), with its own priorities and values, which may compete and conflict with those of humans. Ambiguous terms and concepts, and problematic notions such as ‘orthogonality’, are critically examined, as well as inconsistencies in thought and expectation concerning AGI. Two paradigmatic approaches to the Alignment Problem are compared: Stuart Russell’s and Eric Drexler’s. Large Language Models are discussed in the context of the Turing Test. It is concluded that task-oriented tools, not autonomous agents, should be the goal of AI research.