<p>The ability of highly automated vehicles (HAV) to determine the position of objects in three-dimensional space plays a key role in planning their movement. Implementation of algorithms solving this problem is especially challenging for systems using only monocular cameras, as depth estimation is a non-trivial task in this situation. Nonetheless, such systems are widely used, because of their relatively low cost and ease of operation. We propose here a method for determining the positions of vehicles (the most common type of environmental objects in urban areas) in the form of appropriately oriented bounding boxes in a top view (birds’-eye view) from an image obtained from a single monocular camera. This method consists of two stages. In the first stage, a projection of the visible boundary of the vehicle in the top view is calculated based on 2D obstacle detections and segmentation of the road in the image. It is suggested that the resulting projection represents noisy easurements of two orthogonal sides of the vehicle. In the second stage, an oriented bounding box is constructed around the resulting projection. For this step, we propose a new algorithm for constructing a frame based on the assumption of an L-shaped projection: the L-shape algorithm. The algorithm was tested on a set of real data prepared by ourselves. The proposed L-shape algorithm outperformed the best of the algorithms with which it was compared in terms of the Jaccard coefficient [Intersection over Union, IoU] by 2.7%.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Three-Dimensional Object Detection Based on an L-Shape Model in Autonomous Motion Systems

  • M. O. Chekanov,
  • S. V. Kudriashova,
  • O. S. Shipitko

摘要

The ability of highly automated vehicles (HAV) to determine the position of objects in three-dimensional space plays a key role in planning their movement. Implementation of algorithms solving this problem is especially challenging for systems using only monocular cameras, as depth estimation is a non-trivial task in this situation. Nonetheless, such systems are widely used, because of their relatively low cost and ease of operation. We propose here a method for determining the positions of vehicles (the most common type of environmental objects in urban areas) in the form of appropriately oriented bounding boxes in a top view (birds’-eye view) from an image obtained from a single monocular camera. This method consists of two stages. In the first stage, a projection of the visible boundary of the vehicle in the top view is calculated based on 2D obstacle detections and segmentation of the road in the image. It is suggested that the resulting projection represents noisy easurements of two orthogonal sides of the vehicle. In the second stage, an oriented bounding box is constructed around the resulting projection. For this step, we propose a new algorithm for constructing a frame based on the assumption of an L-shaped projection: the L-shape algorithm. The algorithm was tested on a set of real data prepared by ourselves. The proposed L-shape algorithm outperformed the best of the algorithms with which it was compared in terms of the Jaccard coefficient [Intersection over Union, IoU] by 2.7%.