Drive-CLIP: Cross-Modal Contrastive Safety-Critical Driving Scenario Representation Learning and Zero-Shot Driving Risk Analysis
摘要
Driving risk analysis, especially in safety-critical driving scenarios (SCDSs), plays a paramount role in providing subsequent driving assistance to mitigate potential traffic hazards. Previous studies predominantly relied on unimodal event data for risk analysis, overlooking the valuable source of information embedded in text narratives. This limitation hindered the full utilization of available data for representing SCDSs effectively. Capitalizing on the success of large language models in natural language processing, text-based models offer new potential for enhancing the performance of this task. In this paper, we introduce a novel framework named Drive-CLIP, designed for cross-modal contrastive SCDS representation learning, incorporating both text narratives and event data as two modalities. Through cross-modal analysis, Drive-CLIP distills SCDS event embeddings from natural language supervision, enabling text-guided zero-shot driving risk analysis on event data. Experiments conducted on a naturalistic driving dataset demonstrate that Drive-CLIP surpasses the performance of current best-performing methods, underscoring its effectiveness and superiority. Furthermore, we highlight that cross-modal analysis yields advantages over using a single data modality, and the cross-modal contrastive SCDS representation learning remains beneficial even in scenarios with limited data.