Enhancing Noise Robustness of Speech-Based Human-Robot Interaction in Industry
摘要
In the industrial environment, Human-Robot Interaction is a cornerstone for enhancing productivity and safety. This paper introduces a real-time speech-command recognition (SCR) system designed for running directly on board of embedded devices mounted on board of the robot. Our SCR, based on a ResNet architecture, employs dynamic and domain-specific data augmentation. Our approach significantly enhances accuracy in a noisy environment, achieving a \(+16.2\) improvement over non-augmented training and a \(+3.6\) gain over a non-dynamic approach. These advancements are obtained with low memory usage (1.76 MB) and good processing times (14.9 ms).