A Review of Advances in Large Language and Vision Models for Robotic Manipulation: Techniques, Integrations, and Challenges
摘要
Recent advancements in transformer-based systems, including Large Language Models and Large Vision Models, have significantly transformed robotic manipulation by enabling enhanced task planning, real-time decision-making, and adaptive behaviour in complex environments. This review synthesises current research on integrating these models with robotic control systems, highlighting innovative strategies that merge linguistic and visual processing to improve precision and efficiency. It also critically examines challenges such as scalability, robustness, interpretability, and real-world applicability while identifying research gaps and future directions. This paper provides a concise yet comprehensive overview of the transformative impact of transformer-based systems on robotics, offering valuable insights for developing more sophisticated and versatile robotic systems.