The Elephant in the Room: Why AI Safety Demands Diverse Teams
摘要
Current AI alignment mechanisms are incomplete and may fail as AI systems become superintelligent. We outline a system that could better scale to produce positive outcomes for interaction between humans and superintelligent AI. We consider that existing approaches to AI “safety” and “alignment” may not be using the most effective tools, teams, or approaches. We suggest that an alternative and better approach to the problem may be to treat alignment as a social science problem, since the social sciences enjoy a rich toolkit of models for understanding and aligning motivation and behavior, much of which could be repurposed to problems involving AI models. We use this toolkit to introduce an alternate alignment approach characterized by three facets: 1) defining positive desired social outcomes for human/AI collaboration as the goal or “North Star”, 2) utilizing media archetypes as thought experiments for AI alignment, and 3) forming diverse teams for implementation and oversight of alignment strategies.