错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

IMFA-Stereo: Domain Generalized Stereo Matching via Iterative Multimodal Feature Aggregation Cost Volume

  • Gang Wang,
  • Jinlong Yang,
  • Cheng Wu,
  • Dong Chen

摘要

In the domain of deep learning stereo matching techniques, iterative optimization techniques based on Recurrent All-Pairs Field Transforms (RAFT) have progressively supplanted cost filtering-based methods due to their efficacy in handling extensive scenes. However, they manifest sensitivity to noise and encounter challenges in addressing local blurring issues due to insufficient feature encoding. Consequently, we propose IMFA-Stereo, a hierarchical aggregation network integrating a multi-feature aggregated cost volume. First, we introduce a feature attention-based filtering network to gather multi-class feature information and establish a reliable initial disparity map. Furthermore, we present an aggregated cost volume that encompasses multimodal information, rendering it adaptable for matching tasks across various scenarios. Finally, we employ a ConvGRU-based optimizer to iteratively refine the initial disparity estimation and consolidate the aggregated cost volume, resulting in an accurate disparity map. Comprehensive experiments demonstrate that IMFA-Stereo achieves state-of-the-art stereo matching performance and excels in cross-domain generalization when trained on Scene Flow and applied to real-world scenarios, including KITTI, Middlebury, and ETH3D.