错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LingoQA: Visual Question Answering for Autonomous Driving

  • Ana-Maria Marcu,
  • Long Chen,
  • Jan Hünermann,
  • Alice Karnsund,
  • Benoit Hanotte,
  • Prajwal Chidananda,
  • Saurabh Nair,
  • Vijay Badrinarayanan,
  • Alex Kendall,
  • Jamie Shotton,
  • Elahe Arani,
  • Oleg Sinavski

摘要

We introduce LingoQA, a novel dataset and benchmark for visual question answering in autonomous driving. The dataset contains 28K unique short video scenarios, and 419K annotations. Evaluating state-of-the-art vision-language models on our benchmark shows that their performance is below human capabilities, with GPT-4V responding truthfully to 59.6% of the questions compared to 96.6% for humans. For evaluation, we propose a truthfulness classifier, called Lingo-Judge, that achieves a 0.95 Spearman correlation coefficient to human evaluations, surpassing existing techniques like METEOR, BLEU, CIDEr, and GPT-4. We establish a baseline vision-language model and run extensive ablation studies to understand its performance. We release our dataset and benchmark ( https://github.com/wayveai/LingoQA ) as an evaluation platform for vision-language models in autonomous driving.