Construction method of Wushu tactical decision-making agent driven by knowledge graph and reinforcement learning
摘要
Wushu confrontation processes exhibit strong structural characteristics in terms of tactical organization, continuous action control, and knowledge dependence. Thus, there is a critical need to develop intelligent agents that can comprehensively understand tactical semantics and generate coherent, sequential actions. To solve this problem, this study proposes a Wushu tactical decision-making method based on Knowledge Graph-Guided Hierarchical Reinforcement Learning (KG-HRL), which combines tactical knowledge graph, Graph Attention Network (GAT) and hierarchical strategy structure into the same framework, so that tactical semantics, action correlation and strategy update can simultaneously act on the agent’s tactical decision-making process. In the simulated environment of Wushu confrontation, Knowledge Graph-Guided Hierarchical Reinforcement Learning (KG-HRL) is compared with baseline models such as Deep Q-Network (DQN), Proximal Policy Optimization (PPO), and Asynchronous Advantage Actor-Critic (A3C). The test consists of 1000 independent confrontations. The experiment adopts a strict win-loss determination protocol. If the scores of both sides are the same at the end of a regular round, a secondary determination is made based on the number of effective hits, the number of effective defenses, and the penalty for invalid actions. The test results only include two categories: win and loss. Draws are not included in the win rate. The strict win rate of KG-HRL is 72.3%. The average score is 68.5. The policy stability is 0.82. The action transition cost is 0.12. The tactical diversity index is 12. The control ability of offensive and defensive rhythm is 0.96. The interpretability of knowledge is 80.4%. Compared with the baseline model, KG-HRL demonstrates significant advantages under the same experimental protocol, with a significance level reaching