<p>In the era of exploding video data, efficiently retrieving relevant footage using natural language queries remains a challenge. Existing methods struggle with dynamic backgrounds and lack the ability to capture key moments effectively. Our primary innovation centers on a new method for modeling temporal context, which effectively captures long-term dependencies and contextual information across various time scales. We also integrate a specialized encoder-decoder model that incorporates multi-modal capabilities to extract intricate temporal patterns and interpret user intent from video data. This model effectively addresses temporal variance, enhancing content retrieval and analysis accuracy. We validate our method using extensive comparison analysis on open datasets, with a focus on ranking accuracy using mean average precision for overall ranking quality and recall measures (R@1, R@5) for top highlights. Our findings prove the effectiveness of our framework above current approaches and highlight the advantages of novel temporal context modeling and multi-modal intent.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Query Temporal Context Modeling and Multi-Modal Intent for Efficient Video Content Retrieval

  • Pratibha Singh,
  • Alok Kumar Singh Kushwaha

摘要

In the era of exploding video data, efficiently retrieving relevant footage using natural language queries remains a challenge. Existing methods struggle with dynamic backgrounds and lack the ability to capture key moments effectively. Our primary innovation centers on a new method for modeling temporal context, which effectively captures long-term dependencies and contextual information across various time scales. We also integrate a specialized encoder-decoder model that incorporates multi-modal capabilities to extract intricate temporal patterns and interpret user intent from video data. This model effectively addresses temporal variance, enhancing content retrieval and analysis accuracy. We validate our method using extensive comparison analysis on open datasets, with a focus on ranking accuracy using mean average precision for overall ranking quality and recall measures (R@1, R@5) for top highlights. Our findings prove the effectiveness of our framework above current approaches and highlight the advantages of novel temporal context modeling and multi-modal intent.