错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Scenario-Aware Pareto Optimization for Limited-Round Unknown-Opponent Competition

  • Fanke Chen,
  • JiaKang Mei,
  • Zongyang Liu,
  • Yinghui Pan

摘要

In competitive multi-agent reinforcement learning (MARL), where agents must interact with unknown opponents in a limited number of episodes, traditional reinforcement learning struggles to adapt rapidly, often resulting in suboptimal performance. To address this challenge, we propose a novel Dual-Phase Opponent-aware Architecture (DPOA). This architecture trains a diverse policy set by simulating opponent models with varying strategies, network structures, and training dynamics, thereby producing policies capable of handling a wide spectrum of adversarial behaviors. However, deploying the full set of policies is computationally expensive and impractical in real-time scenarios. To enhance efficiency, we introduce a Scenario-Aware Pareto(SAP) optimization method for policy compression, which selects a compact yet effective subset that balances performance and behavioral diversity across different simulated scenarios. Experiments on the Minefield Navigation Task (MNT) and an Unmanned Aerial Vehicle (UAV) environment demonstrate that our compressed policy set achieves high robustness and adaptability while significantly reducing computational overhead, enabling practical deployment in time-critical settings.