Transformer Based Visual Inertial Odometry
摘要
Visual inertial odometry (VIO) is a sensor fusion technology used for positioning and navigation. It combines visual sensor and inertial sensor information to estimate the movement and location of the UAV in real time. In recent years deep learning based approaches VIO have shown outstanding performance than traditional geometric methods. However, VIO tasks usually need to capture long-distance feature dependencies to ensure the continuity and consistency of camera motion trajectories in time series. In this study, we introduce a new end to end transformer based VIO framework, named VIO-former, to enable the model to better understand motion features in video sequences. Comprehensive quantitative and qualitative evaluation is conducted on KITTI datasets to test our method. The experimental results shows that our approach can achieve superior performance when compared with the existing methods.