Evaluating Vision Language Models (VLMs) across six traditional games: an evaluation paradigm to better assess culturally sensitive question answering systems
摘要
The development of Vision Language Models (VLMs) has made it possible to analyze the cultural dynamics of many cultural art forms from various nations. Robust VLMs have led to the emergence of novel applications of AI (artificial intelligence) and ML (machine learning), including visual question answering (VQA), image captioning, and general scene understanding. However, it is necessary to ensure that VLMs respect a variety of values, prevent ethical misalignments, and promote public trust by addressing accountability, and inclusivity in various cultural contexts. This work focuses on evaluating prominent VLMs in answering cultural queries of six traditional games of Assam, India. Prior works on deep learning-based cultural informatics over traditional games have not explored this dimension. This work proposes a novel evaluation paradigm which considers multidimensional assessment and pair-wise evaluation. It performs cross metric comparison and formulates three algorithms to analyze cultural comprehensibility of two VLMs. A visual question answering (VQA) dataset has been compiled which consists of 3267 image-question pairs for evaluating VLMs through our evaluation paradigm. Qwen2-VL-7B and Gemini 1.5 have been evaluated using three metrics (Cosine similarity and two varieties of LAVE accuracy). This work is going to help one in building an ethical question answering system on top of VLMs.