ReDepthNet: a radar and camera depth estimation model based on semantic segmentation mask region alignment
摘要
This paper proposes a novel depth estimation model, ReDepthNet, based on the alignment of semantic segmentation masks for generating radar depth projections. The model aims to fuse image and radar data to achieve a high-precision depth estimation. A key innovation is the design of a Fusion Encoder that integrates the Efficient Multi-Scale Attention Module (EMA), which efficiently fuses RGB images and radar data. This approach not only captures global features, but also significantly improves the stripe-like artifacts in depth maps and enhances the clarity of object boundaries. Comprehensive experiments conducted on the nuScenes dataset show that, compared to our baseline model, ReDepthNet achieves MAE improvements of 8.2%, 5.5%, and 4.8% at distances of 50 m, 70 m, and 80 m, respectively, and RMSE improvements of 8.2%, 5.4%, and 4.6%, showing significant performance advantages. Furthermore, our experiments highlight the importance of effectively utilizing semantic segmentation mask information and radar point distribution characteristics for depth estimation tasks.