Visual inertial odometry (VIO) is a sensor fusion technology used for positioning and navigation. It combines visual sensor and inertial sensor information to estimate the movement and location of the UAV in real time. In recent years deep learning based approaches VIO have shown outstanding performance than traditional geometric methods. However, VIO tasks usually need to capture long-distance feature dependencies to ensure the continuity and consistency of camera motion trajectories in time series. In this study, we introduce a new end to end transformer based VIO framework, named VIO-former, to enable the model to better understand motion features in video sequences. Comprehensive quantitative and qualitative evaluation is conducted on KITTI datasets to test our method. The experimental results shows that our approach can achieve superior performance when compared with the existing methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Transformer Based Visual Inertial Odometry

  • Sicheng Fei,
  • Jingfeng Li,
  • Lei Li,
  • Jie Liang,
  • Jinwen Hu,
  • Dingwen Zhang,
  • Junwei Han

摘要

Visual inertial odometry (VIO) is a sensor fusion technology used for positioning and navigation. It combines visual sensor and inertial sensor information to estimate the movement and location of the UAV in real time. In recent years deep learning based approaches VIO have shown outstanding performance than traditional geometric methods. However, VIO tasks usually need to capture long-distance feature dependencies to ensure the continuity and consistency of camera motion trajectories in time series. In this study, we introduce a new end to end transformer based VIO framework, named VIO-former, to enable the model to better understand motion features in video sequences. Comprehensive quantitative and qualitative evaluation is conducted on KITTI datasets to test our method. The experimental results shows that our approach can achieve superior performance when compared with the existing methods.