<p>Chemical language models, such as transformers trained on SMILES strings, are increasingly used in drug design and have seen rapid growth in both model capacity and training dataset size. The impact of this scaling on practical downstream performance remains unclear, however. We systematically evaluate how model size and dataset size affect encoder-decoder transformers trained on paired textual molecular representations. We find that, beyond a minimal threshold, further model scaling yields no gain in hit generation rate, while dataset scaling gives diminishing returns. We further introduce a dataset diversification strategy that substantially increases hit diversity. These results suggest that, for molecular hit discovery, data curation and diversity may be more impactful than continued scaling of model size or dataset volume, and they motivate a shift from scale-first to diversity-first training paradigms.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Diversity Beats Size Scaling for Chemical Language Models

  • Borja Medina,
  • Alessandro Tibo,
  • Jiazhen He,
  • Jon Paul Janet,
  • Nicklas Österbacka

摘要

Chemical language models, such as transformers trained on SMILES strings, are increasingly used in drug design and have seen rapid growth in both model capacity and training dataset size. The impact of this scaling on practical downstream performance remains unclear, however. We systematically evaluate how model size and dataset size affect encoder-decoder transformers trained on paired textual molecular representations. We find that, beyond a minimal threshold, further model scaling yields no gain in hit generation rate, while dataset scaling gives diminishing returns. We further introduce a dataset diversification strategy that substantially increases hit diversity. These results suggest that, for molecular hit discovery, data curation and diversity may be more impactful than continued scaling of model size or dataset volume, and they motivate a shift from scale-first to diversity-first training paradigms.