<p>The availability and quality of metadata are key success factors for the effective use of Open Data. However, incomplete or inconsistent metadata significantly hinder the discoverability, interoperability, and reusability of open datasets. This paper examines the extent to which open-source Large Language Models (LLMs) can support the automated generation of high-quality metadata. Based on the Data Catalog Vocabulary (DCAT), a&#xa0;prototype is developed that integrates open-source LLMs into the upload process of Open Data portals. The evaluation using real Open Data datasets demonstrates that the generated metadata often aligns with expert assessments, particularly improving consistency and categorization. However, challenges remain in the precise temporal classification of data and scalability across various data formats. This paper discusses methodological and technical optimization opportunities and provides insights into potential improvements regarding scalability and temporal classification of metadata.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Mit Open-Source-LLMs zu aussagekräftigen Metadaten: Sprachmodelle als Schlüssel für nutzerfreundliche Open-Data-Portale

  • Björn-Lennart Eger,
  • Franziska Ullmann,
  • Barbara Dinter,
  • Peter Gluchowski

摘要

The availability and quality of metadata are key success factors for the effective use of Open Data. However, incomplete or inconsistent metadata significantly hinder the discoverability, interoperability, and reusability of open datasets. This paper examines the extent to which open-source Large Language Models (LLMs) can support the automated generation of high-quality metadata. Based on the Data Catalog Vocabulary (DCAT), a prototype is developed that integrates open-source LLMs into the upload process of Open Data portals. The evaluation using real Open Data datasets demonstrates that the generated metadata often aligns with expert assessments, particularly improving consistency and categorization. However, challenges remain in the precise temporal classification of data and scalability across various data formats. This paper discusses methodological and technical optimization opportunities and provides insights into potential improvements regarding scalability and temporal classification of metadata.