Enhanced Video Event Retrieval Through Adaptive Multi-model Fusion with Large Language Models
摘要
The retrieval of events in videos has emerged as a critical area of research due to the complexity of multimedia information and the rapid growth of digital content. Current methods typically rely on models to extract features from various data sources and combine these features for enhanced retrieval. However, this approach often necessitates prioritizing model weighting based on query context, as different models may yield varying relevance depending on the nature of the query. To address these challenges, this study proposes an adaptive multi-model fusion technique within a video event retrieval framework, leveraging large language models (LLMs) to dynamically adjust weights for multimodal data, thereby enhancing context-aware retrieval. Evaluated on the AI Challenge HCMC (AIC) 2024 dataset, our method achieved a success rate of 91.9%, thereby proving the effectiveness of our solution.