Scenario-Aware Pareto Optimization for Limited-Round Unknown-Opponent Competition
摘要
In competitive multi-agent reinforcement learning (MARL), where agents must interact with unknown opponents in a limited number of episodes, traditional reinforcement learning struggles to adapt rapidly, often resulting in suboptimal performance. To address this challenge, we propose a novel Dual-Phase Opponent-aware Architecture (DPOA). This architecture trains a diverse policy set by simulating opponent models with varying strategies, network structures, and training dynamics, thereby producing policies capable of handling a wide spectrum of adversarial behaviors. However, deploying the full set of policies is computationally expensive and impractical in real-time scenarios. To enhance efficiency, we introduce a Scenario-Aware Pareto(SAP) optimization method for policy compression, which selects a compact yet effective subset that balances performance and behavioral diversity across different simulated scenarios. Experiments on the Minefield Navigation Task (MNT) and an Unmanned Aerial Vehicle (UAV) environment demonstrate that our compressed policy set achieves high robustness and adaptability while significantly reducing computational overhead, enabling practical deployment in time-critical settings.