A Hybrid Video Retrieval System Using CLIP and BEiT-3 for Enhanced Object and Contextual Understanding
摘要
Video retrieval from large datasets has gained significant attention due to its wide range of applications. Existing methods typically focus on either global semantic alignment, which emphasizes detecting prominent objects in queries, or fine-grained semantic alignment, which captures contextual information. However, these approaches often struggle with queries involving complex relationships between objects and their contexts. To address these challenges, we propose a hybrid retrieval approach that integrates Contrastive Language-Image Pre-training (CLIP) and BERT Pre-Training of Image Transformers (BEiT-3). CLIP enhances object detection and recognition, while BEiT-3 excels at understanding detailed contextual relationships. By leveraging the complementary strengths of these models, our approach provides both global semantic understanding and fine-grained contextual analysis across multiple modalities. The proposed system was evaluated in the 2024 Ho Chi Minh City AI Challenge, demonstrating significant improvements in retrieval performance for complex queries.