This paper presents an advanced video retrieval system designed to enhance video content analysis through efficient information extraction and embedding techniques. The system employs a dual approach: videos are indexed by extracting keyframes and reducing redundancies with CNN-based embeddings, using the BEiT-3 model enriched with object detection metadata. Videos are also segmented into overlapping sub-videos, with transcripts aligned and embedded using the Alibaba-NLP model. All processed data is stored in cloud storage and indexed in a vector database for rapid retrieval. User queries are managed via multiple embedding models, enabling versatile searches across transcripts, frames, and descriptions. The retrieval process is refined through a re-ranking algorithm, enhancing relevance and precision. Integrating Retrieval-Augmented methods boosts search accuracy, proving the system’s robustness for large-scale video analysis. The effectiveness of the proposed system is demonstrated through its application in the 2024 AI Challenge, where it successfully addressed complex queries, showcasing its usability and efficiency. This study highlights the potential of the system to transform video retrieval processes, providing a powerful tool for various multimedia applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Video Retrieval via Synergized Image Embeddings and RAG

  • Khac-Toan Nguyen,
  • Vu-Linh Nguyen,
  • Tuyet Hue Tran,
  • Dang-Khoa Nguyen-Le,
  • U. Cao Ky Long,
  • Bui Trong Quy,
  • Tan-Cong Nguyen

摘要

This paper presents an advanced video retrieval system designed to enhance video content analysis through efficient information extraction and embedding techniques. The system employs a dual approach: videos are indexed by extracting keyframes and reducing redundancies with CNN-based embeddings, using the BEiT-3 model enriched with object detection metadata. Videos are also segmented into overlapping sub-videos, with transcripts aligned and embedded using the Alibaba-NLP model. All processed data is stored in cloud storage and indexed in a vector database for rapid retrieval. User queries are managed via multiple embedding models, enabling versatile searches across transcripts, frames, and descriptions. The retrieval process is refined through a re-ranking algorithm, enhancing relevance and precision. Integrating Retrieval-Augmented methods boosts search accuracy, proving the system’s robustness for large-scale video analysis. The effectiveness of the proposed system is demonstrated through its application in the 2024 AI Challenge, where it successfully addressed complex queries, showcasing its usability and efficiency. This study highlights the potential of the system to transform video retrieval processes, providing a powerful tool for various multimedia applications.