Transformers for molecular property prediction: domain adaptation efficiently improves performance
摘要
Over the past six years, molecular transformer models have become a part of the computational toolbox for drug discovery. Most existing models are pre-trained on millions to billions of molecules from large-scale unlabeled datasets such as ZINC or ChEMBL. However, the extent to which such large-scale pre-training improves molecular property prediction remains unclear. This study investigates the potential of transformer models for molecular property prediction while addressing their current limitations. We explore strategies to enhance performance, including the influence of pre-training dataset size and the benefits of domain adaptation through chemically informed objectives. Our results show that increasing the pre-training dataset beyond approximately 400–800K molecules does not improve performance across seven datasets covering five ADME endpoints: lipophilicity, permeability, solubility (two datasets), microsomal stability (two datasets), and plasma protein binding. In contrast, applying domain adaptation on a small number of domain-relevant molecules (
Scientific contribution
We introduce the first systematic evaluation of domain adaptation strategies for molecular transformer models in predicting key ADME properties, such as lipophilicity, solubility, clearance, etc. Our results demonstrate that combining pre-training with domain adaptation significantly improves predictive performance and generalization across diverse chemical datasets (P-value