<p>Universal machine-learning interatomic potentials (uMLIPs) enable near-DFT molecular dynamics (MD) at greatly reduced cost, yet reliability degrades in out-of-distribution (OOD) domains defined by data-sparse chemistries and nonequilibrium configurations. This limitation is particularly consequential for transport properties such as ionic diffusivity, which can be dominated by force errors accumulated along long trajectories. Improving OOD reliability requires targeted data augmentation, yet resource-aware protocols remain scarce, leaving it unclear which structures to label and whether to retrain or fine-tune under a fixed DFT budget. Here, using sodium–oxide solid electrolytes as a representative data-sparse system, we compare three sampling strategies under matched DFT budgets: random selection, DIRECT (uniform coverage in representation space), and LCMD (initial-aware, expanding into regions underrepresented by the foundational dataset). No single strategy is universally optimal. Under scratch training, LCMD is most effective, improving Na-ion diffusivity agreement from R² = 0.48 (MatPES baseline) to R² = 0.70. Under fine-tuning, performance depends primarily on how uniformly the sampled structures cover configuration space, with DIRECT performing slightly better, consistent with its stricter uniform sampling. These results suggest that uMLIP augmentation can be framed as a training-regime-dependent decision problem rather than ad hoc, providing practical guidance for reliable MD-level predictions in multicomponent, data-sparse chemistries.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Budget-constrained augmentation of universal interatomic potentials for reliable multicomponent molecular dynamics

  • Hyeon-Jong Lee,
  • You-Yeob Song,
  • Hojoon Kim,
  • Minjun Kwon,
  • Seung-Hui Ham,
  • Jae-Seung Kim,
  • Donghun Kim,
  • Dong-Hwa Seo

摘要

Universal machine-learning interatomic potentials (uMLIPs) enable near-DFT molecular dynamics (MD) at greatly reduced cost, yet reliability degrades in out-of-distribution (OOD) domains defined by data-sparse chemistries and nonequilibrium configurations. This limitation is particularly consequential for transport properties such as ionic diffusivity, which can be dominated by force errors accumulated along long trajectories. Improving OOD reliability requires targeted data augmentation, yet resource-aware protocols remain scarce, leaving it unclear which structures to label and whether to retrain or fine-tune under a fixed DFT budget. Here, using sodium–oxide solid electrolytes as a representative data-sparse system, we compare three sampling strategies under matched DFT budgets: random selection, DIRECT (uniform coverage in representation space), and LCMD (initial-aware, expanding into regions underrepresented by the foundational dataset). No single strategy is universally optimal. Under scratch training, LCMD is most effective, improving Na-ion diffusivity agreement from R² = 0.48 (MatPES baseline) to R² = 0.70. Under fine-tuning, performance depends primarily on how uniformly the sampled structures cover configuration space, with DIRECT performing slightly better, consistent with its stricter uniform sampling. These results suggest that uMLIP augmentation can be framed as a training-regime-dependent decision problem rather than ad hoc, providing practical guidance for reliable MD-level predictions in multicomponent, data-sparse chemistries.