As vision-language models (VLMs) are increasingly used in various applications, their trustworthiness becomes critical. This paper explores key dimensions of VLM trustworthiness-accuracy, fairness, safety, and robustness. We review methods to enhance visual-textual alignment and tackle challenges like adversarial attacks, biases, and privacy risks through innovative vulnerability assessment frameworks. Our findings highlight ongoing trustworthiness challenges despite some progress. We propose a research agenda emphasizing interdisciplinary approaches to address these gaps and promote responsible VLM deployment.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Trustworthiness in Vision-Language Models

  • Kiana Vu,
  • Phung Lai

摘要

As vision-language models (VLMs) are increasingly used in various applications, their trustworthiness becomes critical. This paper explores key dimensions of VLM trustworthiness-accuracy, fairness, safety, and robustness. We review methods to enhance visual-textual alignment and tackle challenges like adversarial attacks, biases, and privacy risks through innovative vulnerability assessment frameworks. Our findings highlight ongoing trustworthiness challenges despite some progress. We propose a research agenda emphasizing interdisciplinary approaches to address these gaps and promote responsible VLM deployment.