MMMSVR: An Advanced Video Retrieval and Question Answering System
摘要
Video retrieval, the process of locating specific video content within large datasets, presents a significant challenge in the era of digital multimedia. In response to this, as part of the Ho Chi Minh City AI Challenge 2024, this paper presents an advanced multi-modalities and multi-stages video retrieval framework namely MMMSVR, enhanced with the ability to answer supplemental questions, such as automatic counting of objects. The proposed method leverages vision-language models, combining CLIP ViT-H/14, BLIP2 and BEiT-3 for feature encoding and implements a re-ranking mechanism based on a weighting system. Furthermore, a wide range of query modalities such as Optical Character Recognition (OCR), Object Detection, and Automatic Speech Recognition are integrated to refine and improve the retrieval process, supporting both text and image-based queries for efficient retrieval based on multiple attributes. The framework also includes an image-based query feature, enriching the model’s versatility and improving retrieval accuracy. The proposed approach demonstrates significant performance improvements and offers a robust, flexible solution for video search and question-answering tasks.