Multi-scale Spatial-Angular Information Aggregation Network for Image Semantic Segmentation
摘要
Compared to conventional RGB semantic segmentation methods, light field semantic segmentation approaches can incorporate more fine-grained scene information, often leading to superior performance. However, the high-dimensional characteristic of the light field also introduces numerous redundant information and computational burdens. In this paper, we propose a novel light field semantic segmentation network named Light Field Efficient Aggregation Network(LF-EANet), which can efficiently aggregate the structured scene information in light field and generate more precise scene semantic understanding. On the one hand, we encode the spatial position information of light field Sub-Aperture Images(SAIs) to ensure the incorporation of the most valuable perspective position information. On the other hand, we explore multi-scale feature interaction mechanisms from spatial and angular levels to generate more representative auxiliary scene feature. Ultimately, our network provides the center view image with robust scene feature augmentation and semantic perception guidance. This method shows excellent performance on both real-world and synthetic light field semantic segmentation dataset.