Multi-task Perception Model for Unmanned Systems in Urban Environments
摘要
In complex dynamic environments for unmanned systems, achieving efficient and robust multi-task perception is crucial for enhancing environmental understanding and decision-making capabilities. This paper presents a unified multi-task perception framework capable of simultaneously addressing five key perception tasks: depth estimation, pose estimation, optical flow estimation, motion segmentation, and semantic segmentation. The framework employs a shared encoder architecture to improve inter-task synergy through unified feature representations, while integrating both optical flow-based self-supervised depth estimation for dynamic scenes and a Mask2Former-based semantic segmentation model to enhance geometric perception and semantic understanding. By leveraging multi-task collaborative learning, our approach combines spatiotemporal consistency constraints with global semantic information from segmentation to jointly optimize depth estimation and motion segmentation in dynamic scenarios.