DA-MVSNet:depth-aware multi-view stereo network for 3D reconstruction
摘要
We present a novel Depth-Aware Multi-View Stereo (MVS) network, termed DA-MVSNet, for 3D reconstruction from multi-view images. Although existing learning-based methods estimate depth maps effectively through coarse-to-fine strategies, they often neglect important texture and spatial cues in coarse depth maps and fail to fully utilize depth information during cost volume regularization, leading to feature mismatches and suboptimal results. In this paper, we introduce several key innovations to address these limitations. First, we propose a Dual-Branch Interactive Depth Fusion Network (DB-IDFNet) that leverages geometric priors from coarse depth estimates to guide feature extraction in finer stages. Second, we introduce the Haar Wavelet Downsampling (HWD) module in DB-IDFNet to better capture spatial-frequency information, enhancing robustness across different scales. Lastly, we design a Depth-Synergistic Attention Module (DSAM) to capture the importance of varying depth levels and leverage global cross-view information, mitigating mismatches in the cost volume and improving depth estimation accuracy. Experiments conducted on the DTU and Tanks and Temples benchmark datasets demonstrate that our method achieves competitive performance compared to state-of-the-art approaches in terms of depth estimation and 3D reconstruction quality.