错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

GeoClaim: Programmable Geoscientific Fact Verification and Judge-Guided Evaluation for Open-Ended Mineral Exploration QA

  • Yuang Zhang,
  • Pu Zhao,
  • Fanyu Han,
  • Jiaheng Peng,
  • Qinjun Qiu

摘要

Geoscientific and mineral-exploration question answering (QA) requires high factual accuracy, as answers frequently involve spatial topology, coordinate reference systems, and quantitative units. However, widely used automatic metrics such as BLEU and ROUGE rely on surface-level lexical overlap and fail to capture domain-specific factual correctness, resulting in weak alignment with expert judgments for long and professional answers. We propose GeoClaim, a programmable fact-based evaluation framework for open-ended mineral-exploration QA. GeoClaim decomposes both model-generated and reference answers into atomic geoscientific facts (geo-claims) and verifies each claim along three core dimensions: spatial topology, coordinate reference systems and geodesic computations, and unit and numeric equivalence. On this basis, we introduce GeoClaim-F1 and GeoClaim-ROUGE, which measure answer quality at the factual level by assessing claim coverage and consistency, explicitly decoupling semantic correctness from lexical similarity. We further propose Geo-Judge, an evidence-guided LLM-as-Judge mechanism that incorporates structured fact-verification results into the judging process and improves reliability through multi-judge consistency calibration and stability testing. Experiments on a curated dataset of 200 professional mineral-exploration QA pairs show that GeoClaim metrics achieve substantially higher correlation with expert assessments than BLEU and ROUGE, and more accurately reproduce expert model rankings for system comparison and selection. GeoClaim establishes a fact-centric, interpretable, and practically deployable evaluation paradigm for geoscientific QA, supporting reliable assessment and model selection in geological and mineral-resource applications.