Variable Photo-Model Stereo Vision Pose and Size Detection for Home Service Robot
摘要
This paper proposes a method for estimating the pose and size of an object using binocular stereo vision. It utilizes a variable photo-model approach which only requires one shooting condition unknown photograph of the target object. The method constructs a 2D pixel model of the object’s bounding box using the pre-trained YOLOv4 weight from the MS COCO dataset, and converts it into 3D flat photo-models of varying sizes. The estimation of the object’s pose and size is achieved through model-based stereo-vision matching and the use of Genetic Algorithm (GA). This approach eliminates the need for extensive data and pre-training time, making it cost-effective and efficient. Additionally, it extends the application range of the traditional photo-model algorithm. Experimental results conducted in indoor scenes demonstrate the effectiveness of the variable photo-model method in estimating pose and size using a photograph of the same class.