Current AI alignment mechanisms are incomplete and may fail as AI systems become superintelligent. We outline a system that could better scale to produce positive outcomes for interaction between humans and superintelligent AI. We consider that existing approaches to AI “safety” and “alignment” may not be using the most effective tools, teams, or approaches. We suggest that an alternative and better approach to the problem may be to treat alignment as a social science problem, since the social sciences enjoy a rich toolkit of models for understanding and aligning motivation and behavior, much of which could be repurposed to problems involving AI models. We use this toolkit to introduce an alternate alignment approach characterized by three facets: 1) defining positive desired social outcomes for human/AI collaboration as the goal or “North Star”, 2) utilizing media archetypes as thought experiments for AI alignment, and 3) forming diverse teams for implementation and oversight of alignment strategies.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The Elephant in the Room: Why AI Safety Demands Diverse Teams

  • David Rostcheck,
  • Lara Scheibling

摘要

Current AI alignment mechanisms are incomplete and may fail as AI systems become superintelligent. We outline a system that could better scale to produce positive outcomes for interaction between humans and superintelligent AI. We consider that existing approaches to AI “safety” and “alignment” may not be using the most effective tools, teams, or approaches. We suggest that an alternative and better approach to the problem may be to treat alignment as a social science problem, since the social sciences enjoy a rich toolkit of models for understanding and aligning motivation and behavior, much of which could be repurposed to problems involving AI models. We use this toolkit to introduce an alternate alignment approach characterized by three facets: 1) defining positive desired social outcomes for human/AI collaboration as the goal or “North Star”, 2) utilizing media archetypes as thought experiments for AI alignment, and 3) forming diverse teams for implementation and oversight of alignment strategies.