Query Temporal Context Modeling and Multi-Modal Intent for Efficient Video Content Retrieval
摘要
In the era of exploding video data, efficiently retrieving relevant footage using natural language queries remains a challenge. Existing methods struggle with dynamic backgrounds and lack the ability to capture key moments effectively. Our primary innovation centers on a new method for modeling temporal context, which effectively captures long-term dependencies and contextual information across various time scales. We also integrate a specialized encoder-decoder model that incorporates multi-modal capabilities to extract intricate temporal patterns and interpret user intent from video data. This model effectively addresses temporal variance, enhancing content retrieval and analysis accuracy. We validate our method using extensive comparison analysis on open datasets, with a focus on ranking accuracy using mean average precision for overall ranking quality and recall measures (R@1, R@5) for top highlights. Our findings prove the effectiveness of our framework above current approaches and highlight the advantages of novel temporal context modeling and multi-modal intent.