Revising Defeasible Theories via Instructions
摘要
Progress in AI raises agent alignment problems. In this paper, we look at the problem of instructing an agent, i.e. informing it about a regularity in the world it did not previously know. We study an idealized case: agents reasoning with logical theories. The idealization helps to understand the space of possibilities of the problem, and illustrates potential pitfalls and solutions. We believe non-monotonic theories more plausibly approximate human practical and commonsense reasoning so our agents here also use non-monotonic inference. However, instructing a non-monotonic theory does not always result in better alignment. One main cause of this phenomenon is humans omitting the kind of information used by a non-monotonic inference system to resolve conflicts between its parts. We illustrate this with theories induced from a dataset consisting of situated objects. We argue that obtaining non-monotonic theories that respond better to instruction requires additional restrictions on the formalism and theory update procedure.