SDF-Net: Enhanced Novel View Synthesis from Ultra-sparse Viewpoints via Multi-level Feature Fusion
摘要
The primary aim of this study is to enhance the quality of synthetic novel views under conditions characterized by data sparsity. Our investigations demonstrate that the utilization of multi-scale, multi-level stereoscopic depth feature extraction within the cost volume substantially improves the network’s comprehension of scene depth and spatial positioning. Additionally, the introduction of a depth feature beam attention mechanism alleviates the effects of occlusions, thereby enhancing spatial consistency. Specifically, the cost volume is generated through the homography transformation of the feature map, which facilitates deep feature fusion at varying levels and across different spaces using the novel decoder-encoder architecture of the Stereoscopic Deep Fusion Network (SDF-Net). Recognizing the prevalent issue of spatial point occlusion in each beam, we implement an attention mechanism designed to suppress the characteristics of internally occluded spatial points while accentuating those at the extremities of the beam, thus optimizing the effectiveness of spatial features and minimizing occlusion related disturbances during rendering. Extensive experiments show that our approach exhibits state-of-the-art performance compared to previous excellent work on synthesizing novel views when tested on the most popular real scene datasets and synthetic scene datasets, and our results show Richer details and a more complete outline structure.