LameFrames: Optimizing Video Event Retrieval Through Strategic Integration and Individual Strategy Enhancement
摘要
Video event retrieval task aims at retrieving video events from a large video collection that are semantically relevant to a given textual or visual query. There are several approaches have been introduced for this promising task, but overall, they can be categorized either as embedding-based techniques or concept-based techniques. Each of them has its advantages and an appropriately used context. The huge size of datasets often seen in this task is also a challenge for retrieval systems to work efficiently. This paper presents a comprehensive solution and application for video retrieval that addresses the challenges of speed and accuracy in large-scale datasets. Our approach integrates two complementary methods: an image semantic search using CLIP visual-textual embeddings together with an advanced FAISS vector retrieve index; and a concept-based search using optical character recognition (OCR) and automatic speech recognition (ASR) together with Elastic Text Search. In the first approach, we further implement four strategies that advantage and enhance the power of CLIP embeddings. Through experiments, we demonstrate that our approaches provide high accuracy and efficiency for video retrieval applications for a vast type of query. We also found that the quality of the query noticeably impacted the search results and the ability of end users to customize the search by modifying different factors also contributed to a quick and successful query.