<p>The accelerating growth of photographic collections has outpaced manual cataloguing, motivating the use of vision language models (VLMs) to automate metadata generation. This study examines whether AI-generated catalogue descriptions can approximate human-written quality and how generative AI might integrate into cataloguing workflows in archival and museum collections. A VLM (<i>InternVL2</i>) generated catalogue descriptions for photographic prints on labelled cardboard mounts with archaeological content. After a human-in-the-loop curation process, they were evaluated by archive and archaeology experts (consisting of students and professionals) and non-experts in a human-centred, experimental framework. Participants classified descriptions as AI-generated or expert-written, rated quality, and reported willingness to use and trust in AI tools. Classification performance was above chance level, with both groups underestimating their ability to detect AI-generated descriptions. OCR errors and hallucinations limited perceived quality, yet descriptions rated higher in accuracy and usefulness were harder to classify. This suggests that even further human review is necessary to ensure the accuracy and quality of the catalogue descriptions generated by the out-of-the-box model and selected by humans, particularly in specialized domains like archaeological cataloguing. Experts showed lower willingness to adopt AI tools, emphasizing concerns on preservation responsibility over technical performance. These findings advocate for a collaborative approach where AI supports draft generation but remains subordinate to human verification, ensuring alignment with curatorial values (e.g., provenance, transparency). The successful integration of this approach depends on technical advancements, such as domain-specific fine-tuning, and on establishing trust among professionals. Both could be fostered through a transparent and explainable AI pipeline.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ArchiveGPT: A human-centered evaluation of using a vision language model for image cataloguing

  • Line Abele,
  • Gerrit Anders,
  • Tolgahan Aydın,
  • Jürgen Buder,
  • Helen Fischer,
  • Dominik Kimmel,
  • Markus Huff

摘要

The accelerating growth of photographic collections has outpaced manual cataloguing, motivating the use of vision language models (VLMs) to automate metadata generation. This study examines whether AI-generated catalogue descriptions can approximate human-written quality and how generative AI might integrate into cataloguing workflows in archival and museum collections. A VLM (InternVL2) generated catalogue descriptions for photographic prints on labelled cardboard mounts with archaeological content. After a human-in-the-loop curation process, they were evaluated by archive and archaeology experts (consisting of students and professionals) and non-experts in a human-centred, experimental framework. Participants classified descriptions as AI-generated or expert-written, rated quality, and reported willingness to use and trust in AI tools. Classification performance was above chance level, with both groups underestimating their ability to detect AI-generated descriptions. OCR errors and hallucinations limited perceived quality, yet descriptions rated higher in accuracy and usefulness were harder to classify. This suggests that even further human review is necessary to ensure the accuracy and quality of the catalogue descriptions generated by the out-of-the-box model and selected by humans, particularly in specialized domains like archaeological cataloguing. Experts showed lower willingness to adopt AI tools, emphasizing concerns on preservation responsibility over technical performance. These findings advocate for a collaborative approach where AI supports draft generation but remains subordinate to human verification, ensuring alignment with curatorial values (e.g., provenance, transparency). The successful integration of this approach depends on technical advancements, such as domain-specific fine-tuning, and on establishing trust among professionals. Both could be fostered through a transparent and explainable AI pipeline.