A scoping review of reinforcement learning methods for long horizon robotic manipulation
摘要
Long-horizon robotic manipulation presents fundamental challenges for reinforcement learning (RL), including multi-stage decision-making, delayed credit assignment, and the integration of perception, planning, and control over extended time scales. Despite rapid progress, the field remains fragmented across methodological paradigms, task domains, and evaluation practices. This paper presents a data-driven scoping review of 130 studies published between 2020 and 2026, following PRISMA-ScR guidelines, to systematically map the landscape of long-horizon RL for robotic manipulation. We introduce a design-space perspective defined by three structural axes–temporal abstraction, predictive modeling, and structural guidance–and four recurring architectural archetypes–hierarchical, model-based, hybrid, and end-to-end model-free or offline–and provide a descriptive quantitative analysis of task distributions, sensing modalities, action representations, and evaluation protocols. Our analysis reveals a strong concentration of research in simulation-based settings, block rearrangement tasks, and state-based control, alongside limited exploration of real-world-only learning, reset-free execution, dexterous and deformable manipulation, and multimodal perception. To synthesize these findings, we present a gap atlas that identifies underexplored research directions across methodological and experimental dimensions. Finally, we propose a set of standardized reporting recommendations to improve transparency, reproducibility, and comparability in future work. Together, this review provides a comprehensive evidence map of the field and outlines key directions for advancing long-horizon robotic learning toward more realistic and deployable systems.