Coevolutionary Learning in Noncooperative Games
摘要
In this chapter, a case study is presented to demonstrate how principles of competitive coevolution can be applied to develop a multi-agent learning system for the Iterated Prisoner’s Dilemma (IPD) games that are often used to model complex real-world interactions in biological, socio-economic, political, and engineering settings. The first section introduces coevolutionary learning of games in general and highlights how game-playing agents can learn to adapt their decision-making behaviours through coevolutionary process that is only guided by their strategic interactions. Two main coevolutionary learning applications are presented: as a search process for more optimal game-playing strategies and also for simulation purposes to provide answers to what-if scenarios. The next section introduces the background and setting of the problem in this case study. Complex IPD games are presented as models of more realistic real-world interactions. The following section presents our first study on the coevolution of the more complex IPD games involving more choices in agent interactions. Results of our simulated coevolution suggest that the evolution of defection is a result of agents effectively having more opportunities to exploit others when there are more choices. This comes about when agents are less able to resolve the intention of an intermediate choice (e.g. a signal to engender further cooperation or a subtle exploitation). This leads agents to adapt to lower cooperation level plays that offer higher payoffs. We observe that cooperation can occur in complex human interactions that are mediated by mechanism of indirect interactions (e.g. reputation). This motivates the next study we present in the following section on coevolutionary learning of IPD with more choices and reputation. Now, current behavioural interactions depend on not only choices made in previous moves (direct interactions) but also choices made in past interactions that are reflected by their reputation scores (indirect interactions). Our coevolutionary simulations in such settings demonstrate that cooperative behaviours can be learned, whereby agents implement specific strategies that use reputation as a mechanism to estimate behaviours of future partners and to elicit mutual cooperation play right from the start of interactions. We close the chapter with a discussion on implementation issues for which reputation of an agent is estimated (e.g. how frequently reputation scores are updated) and their subsequent impact on the evolution of cooperation.