Learning-based multi-view stereo (MVS) has advanced significantly in inferring depth maps and reconstructing scenes by matching and fusing images from multiple viewpoints, enabling the acquisition of more comprehensive and accurate 3D information. Although significant progress has been made, certain methods still face challenges such as high computational complexity, loss of high-frequency information, and inaccuracies in feature matching. To solve the above problems, we propose WDSNet model, a wavelet transform and depthwise separable convolution based lightweight MVS algorithm designed for 3D reconstruction. The WDSNet model mainly includes the Wavelet Transform-based Feature Pyramid Network (WTFPN) and the Lightweight 3D Harmonize UNet (LHUNet). In WTFPN, we design a feature pyramid network combined with Wavelet Transform, using low-pass and high-pass filters to alleviate high-frequency information loss, effectively preserving important details while reducing the number of learnable parameters. Then, to reduce memory consumption, LHUNet incorporates 3D Depthwise Separable Convolution (3DDS), which decomposes traditional convolution to optimize computational efficiency while minimizing performance degradation. Additionally, 3D Harmonize Attention (3DHA) module is designed to enhance feature matching accuracy by mitigating matching errors across different depths. Experimental results show that our method not only significantly reduces the memory consumption, but also has advantages over other comparison methods in terms of reconstruction results.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Lightweight Multi-view Stereo Method for 3D Reconstruction Using Wavelet Transform and Depthwise Separable Convolution

  • Hu Liang,
  • Bing Liu,
  • Jiacheng Qu,
  • Yuchen Liu,
  • Shengrong Zhao

摘要

Learning-based multi-view stereo (MVS) has advanced significantly in inferring depth maps and reconstructing scenes by matching and fusing images from multiple viewpoints, enabling the acquisition of more comprehensive and accurate 3D information. Although significant progress has been made, certain methods still face challenges such as high computational complexity, loss of high-frequency information, and inaccuracies in feature matching. To solve the above problems, we propose WDSNet model, a wavelet transform and depthwise separable convolution based lightweight MVS algorithm designed for 3D reconstruction. The WDSNet model mainly includes the Wavelet Transform-based Feature Pyramid Network (WTFPN) and the Lightweight 3D Harmonize UNet (LHUNet). In WTFPN, we design a feature pyramid network combined with Wavelet Transform, using low-pass and high-pass filters to alleviate high-frequency information loss, effectively preserving important details while reducing the number of learnable parameters. Then, to reduce memory consumption, LHUNet incorporates 3D Depthwise Separable Convolution (3DDS), which decomposes traditional convolution to optimize computational efficiency while minimizing performance degradation. Additionally, 3D Harmonize Attention (3DHA) module is designed to enhance feature matching accuracy by mitigating matching errors across different depths. Experimental results show that our method not only significantly reduces the memory consumption, but also has advantages over other comparison methods in terms of reconstruction results.