错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Fusion-attention network using dense scale-invariant feature transform flow image and point cloud for 3D pedestrian detection

  • Sang Kyoo Park,
  • Jun Ho Chung,
  • Dong Sung Pae,
  • Tae Koo Kang,
  • Myo Taeg Lim

摘要

In this paper, we introduce a fusion-attention network for three-dimensional (3D) pedestrian detection using the fusion of dense scale-invariant feature transform (SIFT) flow image and point cloud data. Because a pedestrian has a small size and shape, the point cloud data of the pedestrian are insufficient. The absence of point data needs supplementation with an RGB image. However, fusing the RGB image with the point cloud is difficult because of dimensional difference. To fuse the RGB image and point cloud data, point-wise fusion is employed by extracting features from points on the RGB image corresponding to the point cloud. To extract more meaningful features from images, the RGB image is replaced by dense SIFT flow, which represents the movement of the RGB image. To evaluate the proposed method, experimental results were compared with other state-of-the-art methods on the KITTI 3D detection validation set. Then, three ablation studies were conducted. First, to verify the effect of dense SIFT flow, the results were compared with optical flow and RGB image. Second, various fusion methods for dense SIFT flow image and point cloud features were analyzed. Eventually, a final ablation study was set up to ascertain the applicability of the proposed algorithm in a real outdoor driving environment by using a two-dimensional (2D) detection and tracking algorithm instead of a 2D ground-truth boundary box.