Graph-FINDER: multimodal AI for extracting materials data from scientific figures at scale
摘要
A major bottleneck in materials informatics is that quantitative experimental data remain embedded in figures and are inaccessible to machine-readable databases. Here we introduce Graph-FINDER, a multimodal artificial intelligence framework integrating computer vision, natural language processing, and large language models to automatically extract structured numerical data from scientific text and plots. Graph-FINDER digitizes complex multi-line graphs without manual calibration, achieving high shape agreement with curated ground truth (mean Spearman’s ρ > 0.92) and low normalized point-wise error (mean nRMSE = 0.018). Applied to 874 VO2 studies, it reconstructs over 800 resistance–temperature curves, yielding >40,000 datapoints describing transition temperatures, hysteresis widths, and resistance ratios. These data enable application-aware dopant mapping and predictive screening of optimal compositions, with density functional theory supporting AI-identified candidates. Extension to 350 Ge–Sb–Te studies demonstrates cross-material scalability. Graph-FINDER converts fragmented literature into structured datasets, accelerating data-driven materials discovery.