Unstructured environments suffer from a lack of helpful visual patterns for ego-motion estimation, such as the agriculture environments, moreover, intrinsic elements of these environments generate conditions that can be considered difficult for most of the state-of-the-art monocular VO/VSLAM systems, which usually fail to estimate the trajectory and generate a map in this type of environments. Deep learning strategies have been studied to improve the performance of VO/VSLAM systems under diverse approaches in structured environments. Nonetheless, scarce investigations have been oriented to the unstructured ones. In this paper, we address the problem of monocular VO in unstructured agricultural environments. A deep learning-based end-to-end monocular VO system is proposed. Firstly, a study of deep learning-based feature extraction models for monocular VO under the conditions of an unstructured agricultural environment is conducted, for this, state-of-the-art CNN models, which use standard convolution, deformable convolution, and convolutional modulation used in ResNet, InternImage, and Conv2Former blocks as the main operators, were evaluated. Second, inspired by DeepVO system, the best feature extractor model identified in our study, and LTC networks, we build an end-to-end monocular VO system. Evaluation results show that the proposed VO system performs better than baseline DeepVO, even in some sequences competitive to state-of-the-art systems.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Towards End-to-End Visual Odometry for Unstructured Agricultural Environments

  • Víctor Romero-Bautista,
  • Leopoldo Altamirano-Robles,
  • Raquel Díaz-Hernández

摘要

Unstructured environments suffer from a lack of helpful visual patterns for ego-motion estimation, such as the agriculture environments, moreover, intrinsic elements of these environments generate conditions that can be considered difficult for most of the state-of-the-art monocular VO/VSLAM systems, which usually fail to estimate the trajectory and generate a map in this type of environments. Deep learning strategies have been studied to improve the performance of VO/VSLAM systems under diverse approaches in structured environments. Nonetheless, scarce investigations have been oriented to the unstructured ones. In this paper, we address the problem of monocular VO in unstructured agricultural environments. A deep learning-based end-to-end monocular VO system is proposed. Firstly, a study of deep learning-based feature extraction models for monocular VO under the conditions of an unstructured agricultural environment is conducted, for this, state-of-the-art CNN models, which use standard convolution, deformable convolution, and convolutional modulation used in ResNet, InternImage, and Conv2Former blocks as the main operators, were evaluated. Second, inspired by DeepVO system, the best feature extractor model identified in our study, and LTC networks, we build an end-to-end monocular VO system. Evaluation results show that the proposed VO system performs better than baseline DeepVO, even in some sequences competitive to state-of-the-art systems.