Cross-view Transformer for enhanced multi-view 3D reconstruction
摘要
3D reconstruction from multiple 2D images provides rich interactive experiences in design, entertainment, and robotics. However, effectively fusing complementary information across viewpoints remains challenging. This paper proposes a novel cross-view Transformer-based approach for multi-view 3D reconstruction. Our method introduces a cross-view Transformer encoder that achieves effective interaction of information across views. We also develop a global-aware token fusion module to compress multi-view features and an agent attention-based decoder to reduce complexity while maintaining high reconstruction performance. Experiments on benchmark datasets demonstrate that our method achieves state-of-the-art results, outperforming existing techniques.