<p>Analysing large volumes of humanities data requires considerable time and resources due to the high and varied requirements of researchers in terms of data presentation. In addition, humanities scholars wish to enter their search queries in natural language, quickly synchronising distributed data. Efficient approaches, such as DBoD (DataBasing on Demand), can quickly transfer and insert research data into a database. Natural Language Processing (NLP) is used to make these systems more intuitive and user-friendly, and the use of Ontology-based Data Access (OBDA) ensures that the distributed datasets are mapped semantically correctly. In this article, we explain – using a large dataset from the field of epigraphy – how to build a federated information system that supports natural language queries and how OBDA ensures the semantically correct mapping of distributed datasets. We have significantly reduced the implementation time of information systems and a federated information system from months or years to hours. The research data are project-specific and published in a federated information system, complete with the required styles, allowing direct access to the information. In addition, we have a high usability of the federated information systems with NLP-enhanced retrieval.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Ontology-based federated information systems with NLP-enhanced retrieval

  • Sylvia Melzer,
  • Simon Schiff,
  • Franziska Weise,
  • Thomas Asselborn,
  • Meike Klettke,
  • Özgür Lütfü Özçep,
  • Kaja Harter-Uibopuu,
  • Ralf Möller

摘要

Analysing large volumes of humanities data requires considerable time and resources due to the high and varied requirements of researchers in terms of data presentation. In addition, humanities scholars wish to enter their search queries in natural language, quickly synchronising distributed data. Efficient approaches, such as DBoD (DataBasing on Demand), can quickly transfer and insert research data into a database. Natural Language Processing (NLP) is used to make these systems more intuitive and user-friendly, and the use of Ontology-based Data Access (OBDA) ensures that the distributed datasets are mapped semantically correctly. In this article, we explain – using a large dataset from the field of epigraphy – how to build a federated information system that supports natural language queries and how OBDA ensures the semantically correct mapping of distributed datasets. We have significantly reduced the implementation time of information systems and a federated information system from months or years to hours. The research data are project-specific and published in a federated information system, complete with the required styles, allowing direct access to the information. In addition, we have a high usability of the federated information systems with NLP-enhanced retrieval.