The exponential growth of video data on digital media and video-sharing platforms has created an urgent need for efficient content-based video retrieval systems. Traditional methods, such as object recognition, text extraction, and color analysis, have been extensively explored, while current approaches leverage pre-trained multimodal models like CLIP, which give impressive results in text-based video retrieval. However, as the large-scale database gets larger, the level of diversity and complexity increases, and fixed text queries often produce limited results. To address this, we propose an enhanced video retrieval framework that integrates CLIP with GPT-4 through API interaction. By employing advanced prompt engineering techniques, we dynamically expand and refine text queries, enabling broader and more effective exploration of video datasets. Additionally, our framework translates human-generated queries into machine-optimized formats for vision-language models, enhancing retrieval precision. Furthermore, to utilize large image databases such as Google, we introduce an open image-based search functionality that allows users to import reference images similar to their text queries, improving the system’s ability to find relevant content. This dual approach enhances query-model alignment, increasing the likelihood of retrieving accurate and contextually relevant results.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhanced Video Retrieval System: Leveraging GPT-4 for Multimodal Query Expansion and Open Image Search

  • Quang-Khai Bui-Tran,
  • Duc-Huy Ha,
  • Minh-Hung Nguyen,
  • Phuc-Hung Dang,
  • Thien-An Trieu-Hoang,
  • Tien-Huy Nguyen

摘要

The exponential growth of video data on digital media and video-sharing platforms has created an urgent need for efficient content-based video retrieval systems. Traditional methods, such as object recognition, text extraction, and color analysis, have been extensively explored, while current approaches leverage pre-trained multimodal models like CLIP, which give impressive results in text-based video retrieval. However, as the large-scale database gets larger, the level of diversity and complexity increases, and fixed text queries often produce limited results. To address this, we propose an enhanced video retrieval framework that integrates CLIP with GPT-4 through API interaction. By employing advanced prompt engineering techniques, we dynamically expand and refine text queries, enabling broader and more effective exploration of video datasets. Additionally, our framework translates human-generated queries into machine-optimized formats for vision-language models, enhancing retrieval precision. Furthermore, to utilize large image databases such as Google, we introduce an open image-based search functionality that allows users to import reference images similar to their text queries, improving the system’s ability to find relevant content. This dual approach enhances query-model alignment, increasing the likelihood of retrieving accurate and contextually relevant results.