Video Question Answering (VideoQA) remains challenging due to the complexity of video data and the diverse nature of questions. Despite significant advancements in large language models (LLMs) and Vision-language models (VLMs), current VideoQA systems often fall short due to their reliance on shallow reasoning, ineffective frame selection, and a one-size-fits-all approach to answering questions. These limitations hinder their ability to perform deep reasoning, precise localization, and nuanced understanding of causal relationships. In this paper, we propose OptiGQA, a novel framework that integrates the advanced reasoning capabilities of LLMs with the efficient grounding strengths of small Video Grounding Models. OptiGQA enhances VideoQA by rewriting questions to compensate for implicit causal information missed in the query and employing a video grounding model to identify key visual cues. Furthermore, it dynamically selects appropriate processing strategies tailored to different question types, ensuring optimal performance for both global and local queries. Experiments on three standard VideoQA datasets, including NExT-QA, NExT-GQA, and IntentQA, demonstrate that the proposed method outperforms strong baselines and achieves superior localization. Notably, OptiGQA enhances computational efficiency by utilizing fewer frame captions, making it both effective and efficient, advancing the capabilities of VideoQA systems.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

OptiGQA: LLM-Driven Query Optimization for Efficient Visual Grounding in Adaptive Video Question Answering

  • Yufei Wang,
  • Baoyun Peng,
  • Zixuan Dong,
  • Jia Fu,
  • Xinxin Dong,
  • Fei Hu,
  • Xiaodong Wang

摘要

Video Question Answering (VideoQA) remains challenging due to the complexity of video data and the diverse nature of questions. Despite significant advancements in large language models (LLMs) and Vision-language models (VLMs), current VideoQA systems often fall short due to their reliance on shallow reasoning, ineffective frame selection, and a one-size-fits-all approach to answering questions. These limitations hinder their ability to perform deep reasoning, precise localization, and nuanced understanding of causal relationships. In this paper, we propose OptiGQA, a novel framework that integrates the advanced reasoning capabilities of LLMs with the efficient grounding strengths of small Video Grounding Models. OptiGQA enhances VideoQA by rewriting questions to compensate for implicit causal information missed in the query and employing a video grounding model to identify key visual cues. Furthermore, it dynamically selects appropriate processing strategies tailored to different question types, ensuring optimal performance for both global and local queries. Experiments on three standard VideoQA datasets, including NExT-QA, NExT-GQA, and IntentQA, demonstrate that the proposed method outperforms strong baselines and achieves superior localization. Notably, OptiGQA enhances computational efficiency by utilizing fewer frame captions, making it both effective and efficient, advancing the capabilities of VideoQA systems.