<p>In recent years, Vision-Language-Action (VLA) foundation models have been advancing embodied intelligence in robotics by integrating multimodal perception, semantic understanding, and dynamic action generation through end-to-end architectures. This paper focuses on open-source VLA models and their technological innovations and practical applications across three representative robotic domains: Robotic Manipulation, Legged Robots, and Aerial Agents. We systematically analyze their core architectural frameworks, performance advantages, and remaining challenges, providing a comprehensive roadmap for future research and deployment.</p> Graphical Abstract <p></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Open-source vision-language-action models for robotics

  • Linfeng Wang,
  • Deok Jin Lee

摘要

In recent years, Vision-Language-Action (VLA) foundation models have been advancing embodied intelligence in robotics by integrating multimodal perception, semantic understanding, and dynamic action generation through end-to-end architectures. This paper focuses on open-source VLA models and their technological innovations and practical applications across three representative robotic domains: Robotic Manipulation, Legged Robots, and Aerial Agents. We systematically analyze their core architectural frameworks, performance advantages, and remaining challenges, providing a comprehensive roadmap for future research and deployment.

Graphical Abstract