Open-source vision-language-action models for robotics
摘要
In recent years, Vision-Language-Action (VLA) foundation models have been advancing embodied intelligence in robotics by integrating multimodal perception, semantic understanding, and dynamic action generation through end-to-end architectures. This paper focuses on open-source VLA models and their technological innovations and practical applications across three representative robotic domains: Robotic Manipulation, Legged Robots, and Aerial Agents. We systematically analyze their core architectural frameworks, performance advantages, and remaining challenges, providing a comprehensive roadmap for future research and deployment.
Graphical Abstract