Motion saliency-guided multimodal temporal fusion for autonomous driving
摘要
Environmental perception in complex dynamic traffic scenarios is a critical component for the safe operation of autonomous driving systems. To address the limitations of traditional multimodal fusion methods—such as insufficient modeling capabilities for dynamic targets, lack of adaptability in fusion weights, and weak temporal consistency-we propose a motion-saliency-guided multimodal temporal fusion method. This method introduces a motion saliency modeling module that quantifies the dynamic importance of road users by integrating target velocity, trajectory changes, and optical flow information, and uses this to adaptively adjust the fusion weights of multimodal sensors such as cameras, LiDAR, and millimeter-wave radar. Building on this, we construct a fusion framework that combines temporal modeling with consistency constraints to enhance the perception accuracy and robustness of key targets in complex dynamic environments. Experimental results demonstrate that the proposed method outperforms both traditional static fusion methods and typical multimodal fusion methods in scenarios such as highways, urban roads, and complex intersections, achieving superior performance in detection accuracy, temporal stability, and adaptability to dynamic scenes. The results indicate that motion saliency modeling can effectively enhance the representational capabilities of multimodal temporal fusion for high-risk dynamic targets, providing new insights for the design of autonomous driving perception systems.