Information Extraction and Supporting Spatial Constraint Modeling of Plant Habitat Text for Ecological Data Intelligence
摘要
The ecological domain has accumulated massive unstructured textual resources, including floras, species catalogs, and habitat descriptions, which contain rich environmental information such as altitude, hydrology, topography, land cover, and climate. However, expressed in natural language, these texts are often fragmented, compositionally complex, and poorly computable, limiting ecological knowledge organization, data governance, and spatial analysis. To address these issues, this study defines a domain schema with 10 entity types for plant identity and habitat constraints, and proposes a structured extraction framework that integrates cross-category commonality enhancement, difficulty-aware data augmentation, and explicit–implicit joint inference. The framework further alleviates long-tailed category imbalance, overlapping habitat expressions, and implicit ecological constraints using low-frequency sampling, compound habitat template augmentation, overlap-aware hard example mining, and implicit habitat inference. Experiments show that supervised adaptation with data augmentation significantly improves performance. Qwen2-7B with LoRA and augmented data achieves the best Macro-F1 score of 0.9067, far exceeding its zero-shot performance of 0.4020. These results confirm the value of data-centric optimization and task-specific supervision for accurate habitat extraction. The structured outputs can be converted into GIS-compatible ecological predicates, effectively linking habitat text mining to spatially computable ecological constraints.