Large language models for knowledge-centric scientific intelligence: methods, challenges, and lessons from geoscience
摘要
Large language models (LLMs) are increasingly being explored in geoscience, where scientific knowledge is expressed through specialized terminology, heterogeneous documents, maps, imagery, geospatial structures, and temporally ordered interpretations. This review critically examines the emerging literature on geoscience-oriented LLMs (GeoLLMs), focusing on the tasks, construction strategies, evaluation needs, and unresolved challenges that distinguish them from generic LLM applications. We first synthesize the GeoLLM task landscape, including geological information extraction and semantic normalization, relation modeling and knowledge graph construction, evidence-grounded question answering, multimodal map–image–text reasoning, and high-value applications such as hazard-related information analysis, mineral prospectivity evidence synthesis, and chronostratigraphic interpretation. We then review model construction and adaptation strategies, including geoscience corpus engineering, parameter-efficient tuning, domain-adaptive pretraining, retrieval augmentation, ontology and knowledge-graph grounding, and multimodal representation learning. Across these studies, a consistent theme is that GeoLLM outputs should be assessed not only by linguistic fluency, but also by terminology consistency, evidence traceability, spatial and temporal coherence, multimodal grounding, uncertainty expression, and expert validation. The review further identifies major open challenges, including fragmented data and benchmarks, regional and multilingual terminology variation, scale-aware multimodal reasoning, causal and spatiotemporal consistency, hallucination, overtrust, data privacy, proprietary-model dependence, and deployment governance. By using geoscience as a demanding application domain rather than a universal testbed, this review clarifies where domain-specific LLMs can add value, where current evidence remains limited, and what evaluation and governance practices are needed for reliable scientific use.