Video retrieval from large datasets has gained significant attention due to its wide range of applications. Existing methods typically focus on either global semantic alignment, which emphasizes detecting prominent objects in queries, or fine-grained semantic alignment, which captures contextual information. However, these approaches often struggle with queries involving complex relationships between objects and their contexts. To address these challenges, we propose a hybrid retrieval approach that integrates Contrastive Language-Image Pre-training (CLIP) and BERT Pre-Training of Image Transformers (BEiT-3). CLIP enhances object detection and recognition, while BEiT-3 excels at understanding detailed contextual relationships. By leveraging the complementary strengths of these models, our approach provides both global semantic understanding and fine-grained contextual analysis across multiple modalities. The proposed system was evaluated in the 2024 Ho Chi Minh City AI Challenge, demonstrating significant improvements in retrieval performance for complex queries.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Hybrid Video Retrieval System Using CLIP and BEiT-3 for Enhanced Object and Contextual Understanding

  • Long Le Bao,
  • Huy Nguyen Phong,
  • Thang Nguyen Tien,
  • Phat Nguyen Thuan,
  • Thai Hoang Minh,
  • Thuyen Tran Doan,
  • Khiem Le,
  • Tien Do,
  • Duy-Dinh Le,
  • Thanh Duc Ngo

摘要

Video retrieval from large datasets has gained significant attention due to its wide range of applications. Existing methods typically focus on either global semantic alignment, which emphasizes detecting prominent objects in queries, or fine-grained semantic alignment, which captures contextual information. However, these approaches often struggle with queries involving complex relationships between objects and their contexts. To address these challenges, we propose a hybrid retrieval approach that integrates Contrastive Language-Image Pre-training (CLIP) and BERT Pre-Training of Image Transformers (BEiT-3). CLIP enhances object detection and recognition, while BEiT-3 excels at understanding detailed contextual relationships. By leveraging the complementary strengths of these models, our approach provides both global semantic understanding and fine-grained contextual analysis across multiple modalities. The proposed system was evaluated in the 2024 Ho Chi Minh City AI Challenge, demonstrating significant improvements in retrieval performance for complex queries.