Improved weak texture multi-view 3D reconstruction algorithm based on deformable convolutions networks
摘要
DP-MVSNet, a learning-based multi-view stereo (MVS) network utilizing deformable convolutions, addresses the challenge of weak textures in MVS reconstruction. Surpassing the baseline MVSNet architecture, DP-MVSNet introduces deformable convolution modules for enhanced feature extraction and incorporates deformable 3D convolutions into the view weighting network to optimize cross-view informative feature aggregation. DP-MVSNet achieves advanced results on the DTU (0.338 vs. PatchMatchNet’s 0.352) and Tanks & Temples (55.07 vs. PatchMatchNet’s 53.15) datasets, generating denser point clouds with finer details and higher scene completeness. Ablation studies confirm the effectiveness of the deformable convolution-based multiscale feature extraction (DMFE) and deformable 3D convolution-based learning-based patch matching (DLP) modules, reducing overall error by 4.0 % compared to the baseline while maintaining competitive runtime. Demonstrating remarkable robustness in reconstructing realistic landscapes without ground truth calibration, DP-MVSNet represents a significant advancement in accurate and efficient 3D reconstruction of complex environments, particularly those with weakly textured regions.