VizQuest: Enhanced Video Event Retrieval Using Fusion and Temporal Modeling
摘要
Retrieving events from large video datasets is challenging, especially when different data types like images, audio, video, and temporal sequences are involved. Although there have been improvements in cross-modal retrieval, current methods still struggle with precise temporal alignment in event searches. This paper proposes a new video event retrieval system that integrates models for processing text, audio, and visual data. Retrieval accuracy is improved by combining the ranked outputs from these individual models. Each model contributes uniquely to enhance the overall efficiency of the system, and when combined, they enhance both the diversity and robustness of the search. Additionally, temporal modeling ensures the accurate retrieval of event sequences, making the system suitable for time-sensitive tasks. The system was tested in the AI Challenge HCMC 2024, where it demonstrated remarkable speed and precision and secured the top 10, confirming its potential for real-world use in complex event retrieval scenarios.