Multi-task reinforcement learning via mixture-of-patterns meta-learning
摘要
Meta-learning has shown strong performance across reinforcement learning domains, particularly in robotic control and game environments. Yet, rapid adaptation in few-shot robotic learning remains a fundamental challenge. While optimization-based meta-learning methods address this by leveraging meta-priors from task distributions, they face a theoretical limitation: no single model initialization can optimally serve all tasks within complex distributions under homogeneous adaptation protocols. To overcome this, we propose Mixture-of-Patterns Meta-Learning (MoPML), a novel framework that integrates multi-task learning with optimization-based meta-learning. MoPML explicitly decomposes learning into two orthogonal components: discriminative pattern representation via multi-task learning and task optimization via meta-learning. Our approach uses a hierarchical multi-task learner to produce low-dimensional task embeddings that capture the intrinsic structure of the task distribution, while an optimization-based meta-learner utilizes these representations to achieve sample-efficient, task-specific adaptation. Empirical results on a comprehensive set of robotic control benchmarks show that MoPML outperforms state-of-the-art methods in both meta-learning and multi-task learning, with statistically significant gains. Ablation studies and complexity analyses provide insights into the framework's representational capacity and optimization behavior, offering practical guidance for implementation. MoPML not only advances the theoretical foundations of meta-learning but also sets new benchmarks for few-shot adaptation in robotic control.