HW-MLVQA: a novel handwritten multilingual dataset for visual question answering and evaluation
摘要
The proliferation of multilingual Visual Question Answering (VQA) datasets is paramount for augmenting the capabilities of large language models (LLMs) and multi-modal LLMs, thereby enabling them to adeptly capture the intricate linguistic subtleties and visual complexities inherent across diverse languages. This scholarly article delineates HW-MLVQA, a novel dataset meticulously crafted to mitigate the dearth of genuine handwritten datasets pivotal for multilingual document comprehension. HW-MLVQA encompasses an extensive collection of 14,000 handwritten images complemented by 24,000 question-answer pairs. Furthermore, HW-MLVQA provides a robust benchmark evaluation framework spanning three distinct modalities: text, image, and an integrated image & text modality. The dataset also facilitates a rigorous assessment of proprietary and open-source Optical Character Recognition (OCR) systems to simulate a real-world scenario where ground-truth transcriptions are inaccessible. HW-MLVQA aspires to catalyze transformative advancements in multilingual handwritten document interpretation, fostering innovation and scholarly inquiry within this specialized domain. The dataset and code-base will be made publicly available upon acceptance.