NATLOC: Natural Language Object Localization
摘要
This paper presents a novel approach for interacting with a robot in a virtual environment based on one-shot object localization using an image generator model, an image matching model, and a differentially trained robot. The user provides a textual description of an object, which is used by the image generator model to generate a corresponding image. This generated image is then compared with the visual input from a camera mounted on the differential robot, enabling precise object localization when the object is within the camera’s field of view. The robot is trained using Reinforcement Learning techniques to align itself with the requested object. In this way, the robot locates a wide variety of objects solely based on natural language input.