Explainable Strategy Generation Based on Neurally Directed Program Search in Adversarial Environment
摘要
Explainable strategy generation has become a key research focus for robust decision-making in adversarial scenarios. Existing methods, such as imitation learning-based indirect explanations, often suffer from performance loss during transformation, while logical framework-based direct deductions struggle in high-dimensional environments. To address these issues, this study proposes a neural-guided program search method for interpretable strategy generation. The approach uses domain-specific languages (DSL) to preset sketches and employs proximal policy optimization (PPO) to iteratively learn from standard strategies, generating program instructions that form task-specific strategies. Experiments in the Google Football environment show that the proposed method outperforms baselines in winning rate, pass count, and pass efficiency. Ablation studies further validate the necessity of input augmentation and sketches for robust and interpretable strategy generation.