Environmental information is often bound to some geographic entity, be it continent, country, city, or a smaller entity like a forest or a water body. Consequently, documents with environmental context also contain geographic entities. A geo-parser can help to understand what geographic information is present in the document. This information can then be used to display the geographic entities on a map, to group related data as a facet in an internet search, or to enable links between these documents to other documents that refer to the same geographic entity. There have been numerous geo-parsers in the past, however, none of them dealt explicitly with German documents with environmental context. This scenario features a number of challenges that will be explained before a solution is proposed in this publication. As the geo-parser requires some sort of a reference dataset with geographic names and geographic areas, different datasets are analyzed before the most fitting one is picked for the implementation. Furthermore, an evaluation dataset containing pre-tagged geographical entities is created and presented briefly, before the proposed solution is evaluated against the very same dataset.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Geo-Parser for German Documents with Environmental Context

  • Nicolas Doms,
  • Thorsten Schlachter,
  • Lisa Hahn-Woernle

摘要

Environmental information is often bound to some geographic entity, be it continent, country, city, or a smaller entity like a forest or a water body. Consequently, documents with environmental context also contain geographic entities. A geo-parser can help to understand what geographic information is present in the document. This information can then be used to display the geographic entities on a map, to group related data as a facet in an internet search, or to enable links between these documents to other documents that refer to the same geographic entity. There have been numerous geo-parsers in the past, however, none of them dealt explicitly with German documents with environmental context. This scenario features a number of challenges that will be explained before a solution is proposed in this publication. As the geo-parser requires some sort of a reference dataset with geographic names and geographic areas, different datasets are analyzed before the most fitting one is picked for the implementation. Furthermore, an evaluation dataset containing pre-tagged geographical entities is created and presented briefly, before the proposed solution is evaluated against the very same dataset.