Gesture-Controlled Storytelling Framework with Kinematic Actions on ROS and Google TPU-Based Robotic Platform
摘要
Robot storytelling with kinematic actions merges kinematic analysis and narrative techniques, using robotics and artificial intelligence to create engaging and interactive storytelling experiences. This study presents a framework for controlling robot storytelling using human gestures and integrating action kinematics on a Robot Operating System (ROS) and Google tensor processing unit (TPU) based Reachy robot. Utilising Google's MediaPipe model and a deep learning Long Short-Term Memory (LSTM) model for gesture recognition, this responsive storytelling framework pairs with Google Text-to-Speech (TTS) for narrative delivery on the Reachy robot. The methodology involves implementing a gesture recognition model interfaced with ROS, Google PyCoral, and ReachySDK. Eight hand gestures—Open fingers (Start), Close fingers (Stop), Clockwise pointers (Play next), and Anti-clockwise pointers (Play previous)—control storytelling actions with both hands. Real-time gesture recognition is enhanced by the robot's built-in Google Tensor Processing Unit (TPU) for dynamic and fast inferencing. The trained model is optimised for 8-bit quantisation and is compatible with TPU, ensuring adaptability on the ROS platform. Experimentation achieves 98% accuracy in recognising gestures. During storytelling, the Reachy robot utilizes its kinematics and multiprocessing to move its hands, head, and antenna across all 21 degrees of freedom (DoF) for expressive actions. This work contributes to HRI and computer vision by overcoming barriers to speech-based robot control. The gesture-controlled storytelling approach has potential applications in education, entertainment, healthcare, etc. We also present the limitations and future directions of this work.