<p>Over the past six years, molecular transformer models have become a part of the computational toolbox for drug discovery. Most existing models are pre-trained on millions to billions of molecules from large-scale unlabeled datasets such as ZINC or ChEMBL. However, the extent to which such large-scale pre-training improves molecular property prediction remains unclear. This study investigates the potential of transformer models for molecular property prediction while addressing their current limitations. We explore strategies to enhance performance, including the influence of pre-training dataset size and the benefits of domain adaptation through chemically informed objectives. Our results show that increasing the pre-training dataset beyond approximately 400–800K molecules does not improve performance across seven datasets covering five ADME endpoints: lipophilicity, permeability, solubility (two datasets), microsomal stability (two datasets), and plasma protein binding. In contrast, applying domain adaptation on a small number of domain-relevant molecules (<InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\le 4K\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mo>≤</mo> <mn>4</mn> <mi>K</mi> </mrow> </math></EquationSource> </InlineEquation>) using multi-task regression of physicochemical properties significantly improves model performance across all datasets (<i>P</i>-value &lt; 0.001). Furthermore, we find that a model pre-trained on <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(\sim\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>∼</mo> </math></EquationSource> </InlineEquation>400K molecules and adapted on a small domain-specific dataset outperforms larger-scale transformer models like MolFormer and performs comparably to MolBERT. Benchmarking these models alongside baseline representations using RDKit descriptors and Morgan fingerprints reveals that incorporating chemically, physically, and topologically informed features consistently leads to superior performance, regardless of whether used with traditional or transformer-based architectures. While traditional practices such as a random forest model with RDKit descriptors remain strong baselines, this study identifies concrete practices that significantly enhance the performance of transformer models. In particular, aligning pre-training and adaptation with chemically meaningful tasks and domain-relevant data offers a promising path forward for future advancements in molecular property prediction. Our models are available on HuggingFace to allow for easy use and adaptation at <a href="https://huggingface.co/collections/UdS-LSV/domain-adaptation-molecular-transformers-6821e7189ada6b7d0a5b62d4">https://huggingface.co/collections/UdS-LSV/domain-adaptation-molecular-transformers-6821e7189ada6b7d0a5b62d4</a>.</p><p><b>Scientific contribution</b></p><p>We introduce the first systematic evaluation of domain adaptation strategies for molecular transformer models in predicting key ADME properties, such as lipophilicity,&#xa0;solubility, clearance, etc. Our results demonstrate that combining pre-training with domain adaptation significantly improves predictive performance and generalization across diverse chemical datasets (P-value <InlineEquation ID="IEq5"> <EquationSource Format="TEX">\(&lt; 1e^{-9}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mo>&lt;</mo> <mn>1</mn> <msup> <mi>e</mi> <mrow> <mo>-</mo> <mn>9</mn> </mrow> </msup> </mrow> </math></EquationSource> </InlineEquation>). This contribution advances cheminformatics by offering practical guidelines and open resources for developing more accurate and robust molecular&#xa0;transformer models for property prediction.</p> Graphical abstract <p></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Transformers for molecular property prediction: domain adaptation efficiently improves performance

  • Afnan Sultan,
  • Max Rausch-Dupont,
  • Shahrukh Khan,
  • Olga Kalinina,
  • Dietrich Klakow,
  • Andrea Volkamer

摘要

Over the past six years, molecular transformer models have become a part of the computational toolbox for drug discovery. Most existing models are pre-trained on millions to billions of molecules from large-scale unlabeled datasets such as ZINC or ChEMBL. However, the extent to which such large-scale pre-training improves molecular property prediction remains unclear. This study investigates the potential of transformer models for molecular property prediction while addressing their current limitations. We explore strategies to enhance performance, including the influence of pre-training dataset size and the benefits of domain adaptation through chemically informed objectives. Our results show that increasing the pre-training dataset beyond approximately 400–800K molecules does not improve performance across seven datasets covering five ADME endpoints: lipophilicity, permeability, solubility (two datasets), microsomal stability (two datasets), and plasma protein binding. In contrast, applying domain adaptation on a small number of domain-relevant molecules ( \(\le 4K\) 4 K ) using multi-task regression of physicochemical properties significantly improves model performance across all datasets (P-value < 0.001). Furthermore, we find that a model pre-trained on \(\sim\) 400K molecules and adapted on a small domain-specific dataset outperforms larger-scale transformer models like MolFormer and performs comparably to MolBERT. Benchmarking these models alongside baseline representations using RDKit descriptors and Morgan fingerprints reveals that incorporating chemically, physically, and topologically informed features consistently leads to superior performance, regardless of whether used with traditional or transformer-based architectures. While traditional practices such as a random forest model with RDKit descriptors remain strong baselines, this study identifies concrete practices that significantly enhance the performance of transformer models. In particular, aligning pre-training and adaptation with chemically meaningful tasks and domain-relevant data offers a promising path forward for future advancements in molecular property prediction. Our models are available on HuggingFace to allow for easy use and adaptation at https://huggingface.co/collections/UdS-LSV/domain-adaptation-molecular-transformers-6821e7189ada6b7d0a5b62d4.

Scientific contribution

We introduce the first systematic evaluation of domain adaptation strategies for molecular transformer models in predicting key ADME properties, such as lipophilicity, solubility, clearance, etc. Our results demonstrate that combining pre-training with domain adaptation significantly improves predictive performance and generalization across diverse chemical datasets (P-value \(< 1e^{-9}\) < 1 e - 9 ). This contribution advances cheminformatics by offering practical guidelines and open resources for developing more accurate and robust molecular transformer models for property prediction.

Graphical abstract