Much of the research on board games focuses on strong play and on solving games exactly. Recently, programs such as AlphaZero have reached superhuman level in board games such as chess and Go. We study the gap between such AI systems and perfect play, in order to deepen our understanding of their current strengths and limitations. Our study uses Go endgame puzzles with special combinatorial sum game structure, for which an optimal solver is available. We develop an extended Go endgame dataset labelled with exact scores and optimal moves. We evaluate KataGo, the strongest open source AlphaZero-derived program for the game of Go, on these puzzles. We study how the training of different neural networks and the amount of search used affect KataGo’s ability to play perfectly. We observe improved move selection with strong policies, and measure the effect of different MCTS search settings, as well as the challenges KataGo faces in competing against an exact solver. We further analyse move choices in terms of changes of average action value, lower confidence bound, winrate, and number of visited nodes in the MCTS search of KataGo. On our perfect game dataset, KataGo achieves a 90.8% success rate in matches against the exact solver.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Analysing KataGo: A Comparative Evaluation Against Perfect Play in the Game of Go

  • Asmaul Husna,
  • Martin Müller

摘要

Much of the research on board games focuses on strong play and on solving games exactly. Recently, programs such as AlphaZero have reached superhuman level in board games such as chess and Go. We study the gap between such AI systems and perfect play, in order to deepen our understanding of their current strengths and limitations. Our study uses Go endgame puzzles with special combinatorial sum game structure, for which an optimal solver is available. We develop an extended Go endgame dataset labelled with exact scores and optimal moves. We evaluate KataGo, the strongest open source AlphaZero-derived program for the game of Go, on these puzzles. We study how the training of different neural networks and the amount of search used affect KataGo’s ability to play perfectly. We observe improved move selection with strong policies, and measure the effect of different MCTS search settings, as well as the challenges KataGo faces in competing against an exact solver. We further analyse move choices in terms of changes of average action value, lower confidence bound, winrate, and number of visited nodes in the MCTS search of KataGo. On our perfect game dataset, KataGo achieves a 90.8% success rate in matches against the exact solver.