Cheat-FlipIt: An Approach to Modeling and Perception of a Deceptive Opponent
摘要
The modeling of opponent deception in an intelligent game system is not sufficient. However, an opponent agent may launch deceptive actions to consume defense resources, such as feint. We focus on modeling a deceptive opponent. We extend the FlipIt game model and present Cheat-FlipIt model, in which the opponent agent may feint to flip the resources first, and then control the resources after a decay interval. The defense agent models and perceives the cheating behavior of the opponent agent. DQN has some shortcomings such as over-fitting and insufficient exploration, and is not suitable for opponent-deception environment. To address the problems of opponent-deception and non-stationary environment, we present NLD3QN, which incorporates Noisynet, LSTM, Dropout and Dueling Q-Network into the DQN. We further propose a series of cheat strategies of the opponent agent. The defense agent adopts NLD3QN to perceive the cheating behavior of opponents. The proposed approach are evaluated in the Cheat-FlipIt game environments. Experimental results performed on 1 vs. 1 games show that NLD3QN demonstrates superior performance to the baseline DQN. Confronted with a deceptive opponent, the winning rate of NLD3QN is 73.3%, while DQN’s winning rate is 26.67%.