错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Introducing LCC’s NavProc 1.0 Corpus

  • Michael Mohler,
  • Sandra Lee,
  • Mary Brunson,
  • David Bracewell

摘要

In this work, we introduce the NavProc 1.0 Corpus – a medium-scale, annotated corpus of procedural texts within the naval domain – for use as a first step in modeling procedural structures derived from real-world data sources. In particular, we have rigorously produced annotations of frame semantics (i.e., PropBank-inspired trigger/role links) across verbal, nominal, and adjectival frames. Furthermore, we have annotated 21 distinct types of semantic markers and structural links between textual elements (e.g., frame triggers, entities, modifiers) which, taken together, result in a text-focused graph of semantic elements. Such a graph can be used to derive a more complex procedure structure for use in personnel training, simulation, or collaborative procedure execution. Altogether, this annotation effort has encompassed 158 procedural units composed of 2,316 sentences, 44,459 tokens, and 48,137 distinct span annotations. Furthermore, we describe and report LLM-based extraction scores for use as a baseline in future research using this dataset.