Visual Instruction as an Intuitive Interface for Robotic Sorting
摘要
Human-Robot Interaction (HRI) is a rapidly evolving field that seeks to bridge the gap between human users and increasingly capable robotic systems. While robotic technology has advanced significantly, designing intuitive and accessible interaction methods remains a key challenge. Traditional paradigms, such as programming-based interfaces or text-based commands, often impose steep learning curves and hinder broader adoption. While recent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have enabled natural language interaction, these approaches are prone to ambiguity, particularly when specifying spatial relationships. To address these challenges, we propose a sketch-based interaction approach that allows users to communicate tasks and goals through simple, visual inputs, providing a more natural and spatially explicit means of interaction. We apply our method to a robotic sorting task and evaluate its usability and efficiency in a preliminary user study. In this study our sketch-based interface achieved a System Usability Scale (SUS) score of 92.86, indicating excellent usability, and enabled users to instruct the system an average of 5.55 times faster than a text-based approach. These findings suggest that sketch-based interaction can significantly reduce cognitive load and interaction time, making robotic systems more accessible for users of all skill levels. By lowering technological barriers, this approach has the potential to enhance HRI in manufacturing, education, and assistive robotics, paving the way for more seamless and intuitive human-robot collaboration.