Vision-language models (VLMs) represent a transformative advancement in artificial intelligence, combining the processing of textual and visual information to enable sophisticated multimodal applications. The field has witnessed the emergence of diverse architectures and training paradigms consisting of high-capacity models. This research presents a comprehensive evaluation of VLMs in their capability to perform post-processing correction of optical character recognition (OCR) outputs. As we went along, we carefully evaluated the performance of 14 different VLM models, including Mistral, GotOCR, Pali Gemma, and LLaMA 3.2 Vision, in terms of re-captioning of conventional OCR systems like TrOCR, PaddleOCR, EasyOCR, and Tesseract. With a value of 91.22%, Google Paligemma-3B produced the best performance, while Meta Llama-3.2-11 B-Vision came in second with 83.92%. UCASLCL and Deep search VL 1.3 B Base, on the other hand, achieved accuracies of 75.18 and 81.80%, respectively. These developments show several design strategies aimed at improving the model’s performance in general tasks or in particular domains. The performance of Mistral 12B Pixtral and the Adept Fuyu8B, with respective percentages of 59.21 and 9.12% demonstrates the variation of performance across different architectures. This variety also highlights how crucial it is to use standard evaluation techniques in order to create reliable comparisons and direct future advancements in architectural design.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing OCR Post-processing Through Vision-Language Model

  • Fateha Jannat Ayrin,
  • Mahfuzur Rahman Shuvo,
  • Nobanul Hasan,
  • Md Jahid Alam Riad,
  • Prosenjit Roy,
  • Rabeya Nazara,
  • Stabak Das,
  • Vishwanath Akuthota,
  • Md Tanzim Reza,
  • Md Mizanur Rahman

摘要

Vision-language models (VLMs) represent a transformative advancement in artificial intelligence, combining the processing of textual and visual information to enable sophisticated multimodal applications. The field has witnessed the emergence of diverse architectures and training paradigms consisting of high-capacity models. This research presents a comprehensive evaluation of VLMs in their capability to perform post-processing correction of optical character recognition (OCR) outputs. As we went along, we carefully evaluated the performance of 14 different VLM models, including Mistral, GotOCR, Pali Gemma, and LLaMA 3.2 Vision, in terms of re-captioning of conventional OCR systems like TrOCR, PaddleOCR, EasyOCR, and Tesseract. With a value of 91.22%, Google Paligemma-3B produced the best performance, while Meta Llama-3.2-11 B-Vision came in second with 83.92%. UCASLCL and Deep search VL 1.3 B Base, on the other hand, achieved accuracies of 75.18 and 81.80%, respectively. These developments show several design strategies aimed at improving the model’s performance in general tasks or in particular domains. The performance of Mistral 12B Pixtral and the Adept Fuyu8B, with respective percentages of 59.21 and 9.12% demonstrates the variation of performance across different architectures. This variety also highlights how crucial it is to use standard evaluation techniques in order to create reliable comparisons and direct future advancements in architectural design.