<p>The digitization of historical and folkloristic texts presents significant challenges for optical character recognition (OCR), particularly when documents contain complex layouts, embedded illustrations, irregular typography, or non-standard language. This study provides a systematic evaluation of six OCR approaches on two Slovene-language heritage collections: typewritten folklore manuscripts with uniform formatting, and visually heterogeneous issues of the children’s magazine <i>Ciciban</i>. The methods compared include Tesseract, Tesseract with GPT&#xa0;5.2 post-processing, GPT&#xa0;5.2 direct transcription, LLaMA&#xa0;4 Maverick, Nanonets OCR-3, and Qwen-VL-OCR. Performance was assessed using character error rate, word error rate, and complementary sequence-based metrics against manually aligned ground truth. Results indicate that direct multimodal and document-oriented systems achieve the strongest accuracy on typewritten texts, while performance on <i>Ciciban</i> is more sensitive to layout structure. These findings highlight the document-sensitivity of OCR performance and point to the need for adaptive, content-aware pipelines that dynamically integrate multiple OCR strategies. To our knowledge, the study provides the first systematic benchmark of LLM-based OCR for Slovene folkloristic materials, offering practical insights for cultural heritage digitization.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Large language models for OCR in cultural heritage: a comparative study on Slovene Folkloristic texts

  • Octavian M. Machidon,
  • Jasmina Rejec,
  • Domen Vreš,
  • Alina L. Machidon

摘要

The digitization of historical and folkloristic texts presents significant challenges for optical character recognition (OCR), particularly when documents contain complex layouts, embedded illustrations, irregular typography, or non-standard language. This study provides a systematic evaluation of six OCR approaches on two Slovene-language heritage collections: typewritten folklore manuscripts with uniform formatting, and visually heterogeneous issues of the children’s magazine Ciciban. The methods compared include Tesseract, Tesseract with GPT 5.2 post-processing, GPT 5.2 direct transcription, LLaMA 4 Maverick, Nanonets OCR-3, and Qwen-VL-OCR. Performance was assessed using character error rate, word error rate, and complementary sequence-based metrics against manually aligned ground truth. Results indicate that direct multimodal and document-oriented systems achieve the strongest accuracy on typewritten texts, while performance on Ciciban is more sensitive to layout structure. These findings highlight the document-sensitivity of OCR performance and point to the need for adaptive, content-aware pipelines that dynamically integrate multiple OCR strategies. To our knowledge, the study provides the first systematic benchmark of LLM-based OCR for Slovene folkloristic materials, offering practical insights for cultural heritage digitization.