<p>Large language models (LLMs) are increasingly used for information extraction from scientific text, but their reliability in workflows that require structured outputs and executable spatial computation remains uncertain. We evaluate a borehole-report processing pipeline that converts Word documents into structured borehole entities, parses relative location descriptions, resolves survey-control references, and generates start/end coordinates for downstream inspection. Eight representative LLMs were tested on 30 manually annotated reports from a real engineering project under a shared prompt-and-parser protocol, with three repetitions per model-document pair and document-level aggregation. We define coordinate success rate (CSR) as the proportion of location-bearing entities for which the pipeline produces executable coordinates. The results show that high entity and location recall does not guarantee executable coordinate generation: GPT-3.5-Turbo achieved near-perfect recall but CSR&#xa0;=&#xa0;0.188, while the best-performing model reached CSR&#xa0;=&#xa0;0.994. Latency metrics reflect observed provider-route conditions and should not be interpreted as intrinsic model efficiency. We recommend structured-output screening, transparent failure-mode reporting, and cautious interpretation of latency-dependent composite scores when deploying LLM-based geoscience data pipelines. Manually verified coordinate ground truth was not available for this study; CSR therefore measures executable-output compliance rather than geometric accuracy.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LLM-powered borehole data automation for 3D geological modeling workflows: a multi-model evaluation of executable coordinate generation

  • Yuan Liu,
  • Yuanze Du,
  • Yingwang Zhao,
  • Baoping Wang,
  • Hongrui Luo,
  • Xinrui Li,
  • Jinhong Meng,
  • Yuchen Li,
  • Jiahao Ji,
  • Wenxue Tang,
  • Bohan Liu

摘要

Large language models (LLMs) are increasingly used for information extraction from scientific text, but their reliability in workflows that require structured outputs and executable spatial computation remains uncertain. We evaluate a borehole-report processing pipeline that converts Word documents into structured borehole entities, parses relative location descriptions, resolves survey-control references, and generates start/end coordinates for downstream inspection. Eight representative LLMs were tested on 30 manually annotated reports from a real engineering project under a shared prompt-and-parser protocol, with three repetitions per model-document pair and document-level aggregation. We define coordinate success rate (CSR) as the proportion of location-bearing entities for which the pipeline produces executable coordinates. The results show that high entity and location recall does not guarantee executable coordinate generation: GPT-3.5-Turbo achieved near-perfect recall but CSR = 0.188, while the best-performing model reached CSR = 0.994. Latency metrics reflect observed provider-route conditions and should not be interpreted as intrinsic model efficiency. We recommend structured-output screening, transparent failure-mode reporting, and cautious interpretation of latency-dependent composite scores when deploying LLM-based geoscience data pipelines. Manually verified coordinate ground truth was not available for this study; CSR therefore measures executable-output compliance rather than geometric accuracy.