Enhancing Traffic Surveillance with LLaVA: A Multimodal Retrieval Approach for Vehicle Scene Analysis
摘要
The paper introduces an enhanced approach for retrieving vehicle scenes from traffic videos using the fine-tuned Large Language and Vision Assistant (LLaVA) model. Unlike traditional retrieval methods that rely on separate feature extraction for vision and text, our approach integrates vision and language in multimodal learning settings to improve retrieval accuracy. Utilizing the CityFlow-NL dataset from AI City Challenge, which includes extensive traffic video data annotated with natural language descriptions, we applied advanced fine-tuning techniques to improve retrieval accuracy. Our approach achieves a Mean Reciprocal Rank (MRR) of 0.5629, outperforming baseline models such as CNN-based or contrastive learning approaches, demonstrating competitive performance in vehicle scene retrieval. LLaVA enhances vision-text alignment by enabling separate encoders to interact through a learned alignment layer, while fine-tuning further optimizes semantic alignment. These findings underscore the importance of a unified multimodal approach in improving retrieval efficiency for vehicle scene retrieval in intelligent transportation systems.